From 273de4dff3159e463adc110968fea0d956378b14 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 3 Oct 2026 20:46:41 +0000 Subject: [PATCH 01/59] docs: design for ASAP primitive schema and summary semantics One design document for the edge schema of ASAP primitives: fields that carry summary state, the summary semantics a schema must preserve, coverage, and per-operator examples. Related docs keep a summary and link to it. Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/concepts/post-asap-ir.md | 3 +- .../physical-planning-and-deployment.md | 28 +- docs/design_docs/proposals/README.md | 1 + .../proposals/asap-primitive-schema.md | 434 ++++++++++++++++++ .../proposals/decoupling_op_and_expr.md | 2 +- .../design_docs/proposals/operator-sharing.md | 156 +------ .../proposals/univmon-frequency-summary.md | 3 +- .../asap-aware-mapping-contracts.md | 24 +- docs/develop_docs/pre-asap-ir.md | 24 +- 9 files changed, 465 insertions(+), 210 deletions(-) create mode 100644 docs/design_docs/proposals/asap-primitive-schema.md diff --git a/docs/design_docs/concepts/post-asap-ir.md b/docs/design_docs/concepts/post-asap-ir.md index 6c9aa1461..08c670af4 100644 --- a/docs/design_docs/concepts/post-asap-ir.md +++ b/docs/design_docs/concepts/post-asap-ir.md @@ -49,7 +49,8 @@ summary family supports incremental maintenance. and selects the joined rows. Completeness evidence belongs to pruning, not ranking. A `SummaryNode` carries its expression, schema and optional result guarantee. -State and query values have different contracts. Exact operations over +State and query values have different contracts; see +[Schema and physical data for ASAP primitives](../proposals/asap-primitive-schema.md). Exact operations over approximate readouts still require composed accuracy guarantees. See the [accuracy implementation companion](../../develop_docs/end-to-end-accuracy-guarantees.md) and [physical-plan integration](../architecture/physical-plan-integration.md) diff --git a/docs/design_docs/physical-planning-and-deployment.md b/docs/design_docs/physical-planning-and-deployment.md index 274e4974c..c48dfed54 100644 --- a/docs/design_docs/physical-planning-and-deployment.md +++ b/docs/design_docs/physical-planning-and-deployment.md @@ -117,26 +117,14 @@ operators from logical candidates has not completed this integration. ### Input semantics and summary semantics -`source`, `filter`, `grouping` and `window` describe input-data semantics: -where records originate, which records qualify, how they are grouped and which -time interval applies. They are not a complete description of arbitrary summary -computation. In particular, the same four fields can summarize different value -expressions or produce different states. - -| Concern | Required semantic information | -| --- | --- | -| Input computation | Source identities and schemas, filters, joins/transforms and their order, or a reference to the canonical input sub-DAG | -| Values and grouping | Value expressions, item identities and weights where applicable, group keys and types, and operation-defined null/duplicate handling | -| Time | Time column and interpretation, interval bounds, evaluation alignment, and distinction between query range and maintained panes | -| Summary computation | Exact operation or sketch family, algorithm and parameters, and supported build/merge behavior | -| Output | State versus finalized value, output schema/type, and readout parameters when part of the output computation | - -For example, KLL over `latency_seconds` and KLL over `log(latency_seconds)` differ -even with identical source, filter, grouping and window. Likewise, weighted -frequency state needs both item and weight expressions. More complex inputs -must retain their computation DAG; four descriptive fields cannot replace it. - -The canonical selected computation is authoritative. These categories describe +`source`, `filter`, `grouping` and `window` describe input-data semantics but not +a complete summary computation: the same four fields can summarize different +value expressions or produce different states. The semantic information a summary +depends on, and where the IR records each part (field type, producing operator, +or coverage), is specified in +[Schema and physical data for ASAP primitives](proposals/asap-primitive-schema.md#23-consideration-3-the-metadata-preserves-summary-semantics). + +The canonical selected computation is authoritative. Those categories describe what must be preserved, not a new flat IR or a second expression language. Operator-defined behavior should be referenced through its canonical contract, not independently configured in deployment metadata. Unsupported or unresolved diff --git a/docs/design_docs/proposals/README.md b/docs/design_docs/proposals/README.md index cb44fadbb..6c4ed49f8 100644 --- a/docs/design_docs/proposals/README.md +++ b/docs/design_docs/proposals/README.md @@ -11,3 +11,4 @@ extensions. A design document is not a promise of downstream runtime support. - [Operator sharing](operator-sharing.md) - [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) - [ASAPPlanner layering](planner-layering.md) +- [Schema and physical data for ASAP primitives](asap-primitive-schema.md) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md new file mode 100644 index 000000000..8840a892d --- /dev/null +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -0,0 +1,434 @@ +# Schema and Physical Data for ASAP Primitives + +> Status: design. `SummaryCoverage` and its rules are implemented in [#567](https://github.com/ProjectASAP/ASAPPlanner/pull/567); `SummaryMerge` in [#560](https://github.com/ProjectASAP/ASAPPlanner/pull/560) (both open). Audience: designers and architects. + +This document is the single source of truth for the schema, and column design for ASAP Primitives. This is used in the logical stage (LogicalASAPDAG), and physical stage (PhysicalASAPDAG). + +## 1. Goal, problem, and requirements + +Unlike existing Database engines, which work on raw data or explicitly defined materialized tables with schema and column names provided by the users, ASAPPlanner is designed for querying and execution over the mix of raw data and ASAP Primitives. ASAP primitives are usually compact summaries over raw data. Therefore, it introduces new requirement when we design the schema and node definitions for LogicalASAPDAG and PhysicalASAPDAG. + +Assuming we have the Logical DAG defined for a canonicalized representation for a batch of queries. [TODO: add links for this here. ] +The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG [TODO: link PR 511's doc here], but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: `SummaryAgg` (creates and updates state; there is no separate SummaryCreation or SummaryUpdate operator, and `SummaryUpdate` is `SummaryAgg`'s update-expression parameter), `SummaryEstimate`, `FinalizeExactAccumulator`, `MaintainPopulation` and `EvaluatePopulation` (implemented); `SummaryMerge` (reserved here, enabled by #560); and `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` (reserved). See §5. +Each of the Summary operators also require the ASAP primitive information above to inter-operate correctly, preserving semantic correctness. + +Basically, the following information should be represented to preserve the equivalent query semantics when we introduce ASAP Primitives to logical query representation, and following physical one. + +- What type of the ASAP Primitive is +- What is the ASAP Primitive parameters +- What data sources a ASAP primitive summarizes +- What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents + + + +And these information will be combined with relational or time series query operator information, such as group by/reduction, filtering, projection, join, time series selection, together. + +Therefore, these requirements drive the following schema and metadata, node information, and column design. + + + +## 2. Existing database terminology for schema, table, column, and physical data layout + + +## 3. Proposed schema design +Schema represents the **metadata** of information flow along an **edge** between two nodes in a logical or physical DAG. The schema field is associated with the node in the DAG. The consumer of the node in the DAG takes the schema from the producer node as input. + +Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. +Based on our requirement, each field should contain the following information. +1. **What type of the ASAP Primitive is** A state column can be a raw data type (e.g., numerical number, string). It can also be a summary type (§6.1, §6.2), e.g., the summary **family** is sketch (`FieldDataType::Sketch`), its **category** is quantile (`SketchCategory::Quantile`), its **algorithm** is KLL (`SketchAlgorithm::Kll`), and its **parameters** are `SketchParams::Kll { k: 200 }`. A sketch field also records its `GroupingStrategy` (one instance per group, or one shared Hydra structure). Non-sketch families (`ExactAggregate`, `Sample`, `Wavelet`, `StatModel`) have a kind and parameters but no category. +2. **What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents** This information is being mapped based on the primitive type: it is not stored in the field. `SummaryEstimate::validate_inputs` accepts a readout only when the `SketchStatistic` matches the state's `SketchCategory` (Quantile→`Quantile`; Cardinality→`Cardinality`/`Universal`; PointCount→`Frequency`/`Universal`; FrequencyL2/FrequencyEntropy→`Universal`; TopK→`TopK`/`Universal`). Exact accumulators are read by `FinalizeExactAccumulator` instead. `Sample`, `Wavelet` and `StatModel` state has no readout operator yet. + +Whether an edge carries **state or values** is recorded on the producing node as `OperatorResultKind`, not inferred from field types. Usually they agree: a `SummaryAgg` output has exactly one non-plain field and kind `State`. They can differ, though. `MaintainPopulation` keeps an all-plain schema but its kind is `State`, and a `Project` that passes a state column through keeps kind `State`. Consumers check `result_kind` (`validate_inputs` requires `State` for every readout). + +## 4. Proposed Node field design + +A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. + +In code this field is `OperatorNode::coverage`. It records *which observations* a summary state covers: a source plus a union of joint (time × population) regions. It sits beside the schema, not inside a field, because two states with identical schemas can cover different data, and only disjoint coverage can be merged once-per-observation. + +Based on the above the proposed OperatorNode interface is as below (`crates/types/src/ir/node.rs`, `ir/summary_coverage.rs`): +```rust +pub enum OperatorResultKind { Relation, InstantVector, RangeVector, /** unfinalized summary/accumulator state */ State } + +pub enum Operator { NonASAP(NonASAPOp), ASAP(ASAPOp) } + +pub struct OperatorNode { + pub operator: Operator, + pub result_kind: OperatorResultKind, // derived from operator + children + pub schema: Schema, // derived; may override names/qualifiers only + pub guarantee: Option, // None until accuracy assessment; None != exact + pub timing: Option, // IngestionTime | QueryTime; None until assigned + #[serde(default)] + pub coverage: Option, // observations a State output summarizes +} + +impl OperatorNode { + pub fn new(op: Operator) -> Result; // derives schema + kind; coverage = None + pub fn with_schema(op: Operator, schema: Schema) -> Self; // caller-supplied names + pub fn new_shared(op: Operator) -> Result, SchemaDerivationError>; + pub fn with_guarantee(self, g: Option) -> Self; + pub fn with_timing(self, t: Option) -> Self; + /// Validates the coverage, then rejects a non-State node (CoverageError::NotState). + pub fn with_coverage(self, c: SummaryCoverage) -> Result; + /// This branch: true only for SummaryAgg. #560: SummaryAgg | SummaryMerge. + pub fn requires_coverage(&self) -> bool; + /// #560: (update expression, reduction) of a SummaryAgg, or shared by a SummaryMerge's inputs. + pub fn summary_update(&self) -> Option<(&SummaryUpdate, &Reduction)>; + /// Rebuilds with new inputs; re-derives schema; clears guarantee, timing and coverage. + pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Result; + /// Per node: coverage well-formed (and present if required), validate_inputs, + /// result_kind and schema structure agree with derivation. + pub fn validate_structure(self: &Rc) -> Result<(), SchemaDerivationError>; + pub fn validate_execution_timing(self: &Rc) -> Result<(), SchemaDerivationError>; + // also: asap(), non_asap(), is_asap(), children(), contains_asap(), reachable() +} + +pub struct SummaryCoverage { + pub source: Source, // Source::Table { table_ref } | Source::TimeSeries { metric } + pub regions: Vec, // union of joint regions, never a Cartesian product +} +pub struct CoverageRegion { + pub time_ms: Option>, // half-open, on the source's time column; None = unrestricted + pub population: BTreeMap, // conjunction of equality predicates; empty = unrestricted +} +impl SummaryCoverage { + /// Intervals non-empty, dimension names non-empty, regions pairwise provably disjoint. + pub fn validate(&self) -> Result<(), CoverageError>; + /// Union of provably disjoint inputs from one source; coalesces adjacent + /// intervals with identical populations and keeps gaps. + pub fn merge_disjoint(inputs: &[Self]) -> Result; +} +pub enum CoverageError { + InvalidInterval, InvalidPopulation, SourceMismatch, PossibleOverlap, EmptyMerge, NotState, Missing, + UnknownInput, // #560: a merge input has no coverage + MergeOutputMismatch, // #560: retained merge coverage differs from the input union +} +``` + +Two regions are provably disjoint only if their time ranges do not intersect or they bind the same population dimension to different values. A region without time bounds overlaps any region it is not population-disjoint from. Coverage is caller-established: `validate` checks that it is well-formed, not that it matches the child's predicates. Any rewrite through `map_children` drops it, so the rewriter must declare it again. `#537` adds export and CSE of the logical DAG (`ir/export.rs`, `ir/cse.rs`); no interface in this document depends on it. + +## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators + +Given that these information requirements are introduced by summary operators to work correctly semantically, we show the examples of how the defined OperatorNode, schema, and physical data information work with each kind of summary operators. + +Notation: an edge is written `──Kind(field Type, …)──▶`. Schemas are the ones `output_schema()` derives. Planning may rename fields through `OperatorNode::with_schema`, but types, nullability, `time_index`, `unique_keys` and `closed` must match the derivation. All examples use a table source, so values are `Relation`; with a `TimeSeries` source the value side is `InstantVector`. + +### 5.1 `SummaryAgg`: values → state + +Scenario: p99 latency by job, from KLL(k=200), over one minute of table `t`. + +```text +Scan(t: job Utf8, latency Float64) + ──Relation(job Utf8, latency Float64)──▶ +SummaryAgg(family = Sketch(KLL{k=200}, PerSubpopulationInstance), + input = SummaryUpdate::column(Named("latency")), reduction = by[job], + grouping = PerSubpopulationInstance, filter = None) + ──State(job Utf8, state Sketch(KLL{k=200}, PerSubpopulationInstance))──▶ + coverage = { source: Table "t", regions: [{ time_ms: 0..60_000, population: {} }] } +``` + +- Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. +- Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. +- Coverage: **required** and **declared**. `OperatorNode::new` leaves it `None`, `validate_structure` fails with `CoverageError::Missing`, and the planner attaches it with `with_coverage`. +- Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. + +### 5.2 `SummaryEstimate`: sketch state → value + +Scenario: read p99 from the state in 5.1. + +```text +──State(job Utf8, state Sketch(KLL{k=200}))──▶ +SummaryEstimate(query = SketchStatistic::Quantile { q: 0.99 }) + ──Relation(job Utf8, quantile Float64)──▶ (planner may rename to p99) +``` + +- Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`), and `topk` Utf8. Keys and metadata pass through. +- Result kind: the value kind of the source the state was built from (`Relation` here). +- Checks: input is `State` with exactly one non-plain field, that field is `Sketch`, and its category accepts the statistic (§3). For example, `Cardinality` on KLL is rejected. +- Coverage: **absent**. The output is a value, and `with_coverage` returns `NotState`. +- Boundary: state is consumed and a value is produced; `guarantee` on this node carries the readout's error bound. + +### 5.3 `FinalizeExactAccumulator`: exact state → value + +Scenario: total bytes by host with an exact Sum accumulator. + +```text +Scan(t: host Utf8, bytes Float64) + ──Relation(host Utf8, bytes Float64)──▶ +SummaryAgg(family = ExactAggregate(Sum, Sum), input = column(Named("bytes")), reduction = by[host]) + ──State(host Utf8, state ExactAggregate(Sum, Sum))──▶ coverage: required, declared +FinalizeExactAccumulator + ──Relation(host Utf8, state Float64)──▶ +``` + +- Output schema: each `ExactAggregate` field keeps its name (`state`) and takes the type and nullability the equivalent `NonASAPOp::Aggregate` would give: Sum/Min/Max follow the input column, Count is Int64, and Rate/IRate/Increase are Float64. If the child is not a `SummaryAgg` directly, Count falls back to Int64 and the others to Float64. `unique_keys`, `closed` and `time_index` are preserved (`schema_rebuilding.rs`). +- Checks: the input is `State` and contains an `ExactAggregate` field; a sketch is rejected (`structure_contract.rs`). +- Coverage: **absent** on the output. +- Boundary: this is the explicit maintenance-to-read boundary for exact state. Exact state is never read through `SummaryEstimate`. + +### 5.4 `MaintainPopulation`: values → maintained membership (state) + +Scenario: keep the full latency population per job, so that p99 and top-10 can be evaluated later. + +```text +Scan(t: job Utf8, latency Float64) [closed schema] + ──Relation(job Utf8, latency Float64)──▶ +MaintainPopulation(population = MaintainedPopulation { + input: PopulationInput::Rows { input: , value_column: 1, grouping: by[job] }, + max_k: 10, quantiles: true }) + ──State(job Utf8, latency Float64)──▶ +``` + +- Output schema: identical to the child's, all plain. Only `result_kind = State` marks it as maintained state. +- Checks: `population.matches_node(child)`. For `Rows`, the child must be the same closed table `Scan`, the value column must be non-null Float64, and grouping must be `by` with in-range keys. For `CurrentSeries`, it must be a `TimeSeries` scan with the same metric, matchers and grouping labels, under an instant `TimeRange` of `lookback_ms` (which may be omitted only for the default 300 s lookback). +- Coverage: **not required**. `with_coverage` accepts it because the output is `State`. +- Boundary: the output is state because it must also track membership changes; downstream operators can only read it through `EvaluatePopulation`. + +### 5.5 `EvaluatePopulation`: maintained membership → value + +Scenario: p99 by job from the population in 5.4. + +```text +──State(job Utf8, latency Float64) [from MaintainPopulation]──▶ +EvaluatePopulation(evaluation = PopulationStatistic::Quantile { q: 0.99 }) + ──Relation(job Utf8, quantile_0_99 Float64)──▶ +``` + +- Output schema: the schema of `Aggregate(by grouping, measure)` over the maintained source. Quantile gives `quantile_` Float64, Sum gives `sum` (value type), Count gives `count` Int64 and Average gives `avg` Float64; `unique_keys = [[0]]`, `closed`. `TopK { k }` instead returns the source schema unchanged (the selected rows). +- Checks: the child is a `MaintainPopulation` node whose `supports(evaluation)` holds: `quantiles` must be set for `Quantile`, and `k <= max_k` for `TopK`. +- Coverage: **absent**. +- Boundary: maintained membership is read as a value; the result kind is the source's (`Relation`). + +### 5.6 `SummaryMerge` (#560): state × N → state + +On this branch, `SummaryMerge { children }` is **reserved**. `is_unimplemented()` returns true, and `output_schema()` and `validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so `OperatorNode::new` fails. Only `output_kind()` (= `State`), `children`, `map_children` and `kind_name` work. #560 enables it as follows (`summary_merge_structure.rs` in #560). + +Scenario: combine two one-minute KLL panes over `Scan(t: value Float64)` into a two-minute state. + +```text +SummaryAgg(KLL k=200, column(SampleValue), by[]) ──State(state Sketch(KLL{k=200}))── coverage {t, [0..60_000)} ─┐ +SummaryAgg(KLL k=200, column(SampleValue), by[]) ──State(state Sketch(KLL{k=200}))── coverage {t, [60_000..120_000)} ─┴▶ +SummaryMerge + ──State(state Sketch(KLL{k=200}))──▶ coverage = { t, [0..120_000) } (derived) +``` + +- Output schema: `children[0].schema`. +- Checks: at least one input; exactly one state column; every input is `State` with an identical schema (so family, params, grouping strategy and key positions match); every input has the same `summary_update()` (update expression and reduction); and `merged_coverage()` succeeds. Merging k=200 with k=300 fails, and so does merging raw rows. +- Coverage: **required** and **derived**. `OperatorNode::new` sets it to `SummaryCoverage::merge_disjoint` of the input coverages. An input without coverage gives `UnknownInput`, and overlapping inputs give `PossibleOverlap`. `validate_structure` rejects a retained coverage that differs from the derived one (`MergeOutputMismatch`). Gapped inputs stay as two regions. +- Boundary: state in, state out. No value is produced until a readout. + +### 5.7 Reserved operators (not implemented) + +These variants exist so that plans can name them, but `output_schema()`/`validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so no node can be built. `output_kind()` already returns `State` for each of them. The intended edge shapes below follow from their fields; none of them is implemented. + +| Operator | Fields | Intended edge shape | +|---|---|---| +| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5) | +| `SummaryDelete` | `summary_input, key: ColumnId` | State → State with the entries for `key` removed | +| `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | +| `Extension` | `child, name` | deployment-named state operator | + +### 5.8 Summary + +| Operator | Input kind | Output kind | Output carries state | Coverage on output | Status | +|---|---|---|---|---|---| +| `SummaryAgg` | value (not `State`) | `State` | yes (one `family` field) | required, declared | implemented | +| `SummaryEstimate` | `State` (one `Sketch` field) | source's value kind | no | absent | implemented | +| `FinalizeExactAccumulator` | `State` (`ExactAggregate`) | source's value kind | no | absent | implemented | +| `MaintainPopulation` | `Relation` (table) / `InstantVector` (series) | `State` | yes (by kind; fields plain) | optional, not required | implemented | +| `EvaluatePopulation` | `State` from `MaintainPopulation` | source's value kind | no | absent | implemented | +| `SummaryMerge` | `State` × N | `State` | yes | required, derived | reserved; enabled by #560 | +| `SummarySubtract` | `State` × 2 | `State` | yes | — | reserved | +| `SummaryDelete` | `State` | `State` | yes | — | reserved | +| `SummaryJoin` | `State` × 2 | `State` | yes | — | reserved | +| `Extension` | any | `State` | yes | — | reserved | + +## 6. Key code interfaces + +`OperatorNode`, `OperatorResultKind` and coverage are in §4. Bodies and serde/derive attributes are elided below. + +### 6.1 Schema and field types (`crates/types/src/pre_asap/schema.rs`) + +```rust +pub type ColumnId = usize; + +pub struct Schema { + pub fields: Vec, + pub time_index: Option, // must point at a plain Timestamp field + pub unique_keys: Vec>, + pub closed: bool, // true = fields enumerate every column +} +impl Schema { + pub fn new(fields: Vec) -> Self; + pub fn with_time_index(fields: Vec, time_index: ColumnId, unique_keys: Vec>) -> Self; + pub fn lifted(fields: Vec, time_index: Option) -> Self; // closed = true + pub fn is_all_plain(&self) -> bool; + pub fn column_id(&self, name: &str) -> Option; + pub fn column_id_qualified(&self, table: &str, name: &str) -> Option; +} + +pub struct Field { + pub name: String, + pub dtype: T, + pub nullable: bool, + pub table: Option, +} +impl Field { + pub fn plain(name: impl Into, dtype: DataType, nullable: bool) -> Self; + pub fn plain_dtype(&self) -> Option<&DataType>; + pub fn is_plain(&self) -> bool; +} + +/// A column's type: a plain value, or summary state of one family. +pub enum FieldDataType { + Plain(DataType), + ExactAggregate(ExactKind, ExactParams), + Sketch(SketchKind, GroupingStrategy), + Sample(SamplingKind, SamplingParams), + Wavelet(WaveletKind, WaveletParams), + StatModel(StatModelKind, StatModelParams), +} + +pub enum DataType { + Null, Int64, Float64, Utf8, Bool, Timestamp, Interval, Date, + List { element: Box> }, + Struct { fields: Vec> }, + Map { key: Box, value: Box, value_nullable: bool }, +} +``` + +### 6.2 State-family parameters (`crates/types/src/post_asap/sketch.rs`) + +```rust +pub enum ExactKind { Sum, Count, Min, Max, Increase, Rate, IRate } +pub enum ExactParams { Sum, Count, Min, Max, Increase, Rate, IRate } // no knobs; mirrors kind + +pub struct SketchKind { category: SketchCategory, algorithm: SketchAlgorithm, params: SketchParams } +impl SketchKind { + /// The only constructor; classifies the category and panics on mismatched params. + pub fn new(algorithm: SketchAlgorithm, params: SketchParams) -> Self; + pub fn category(&self) -> SketchCategory; + pub fn algorithm(&self) -> &SketchAlgorithm; + pub fn params(&self) -> &SketchParams; +} +pub enum SketchCategory { Universal, Quantile, Cardinality, Frequency, TopK } +// Universal: UnivMon | Quantile: Kll, DDSketch | Cardinality: Hll, Theta, Kmv +// Frequency: Cms, CountSketch | TopK: CmsWithHeap, CountSketchWithHeap +pub enum SketchAlgorithm { UnivMon, Kll, Cms, Hll, DDSketch, CmsWithHeap, Kmv, Theta, CountSketch, CountSketchWithHeap } +pub enum SketchParams { + UnivMon { heap_size: u32, sketch_rows: u32, sketch_cols: u32, layers: u8 }, + Kll { k: u32 }, + Cms { width: u32, depth: u32 }, + Hll { precision: u8 }, + DDSketch { alpha: f64 }, + CmsWithHeap { width: u32, depth: u32, heap_size: u32 }, + Kmv { k: u32 }, + Theta { k: u32 }, + CountSketch { width: u32, depth: u32 }, + CountSketchWithHeap { width: u32, depth: u32, heap_size: u32 }, +} + +/// How grouped state is instantiated across `by` subpopulations. Orthogonal to family. +pub enum GroupingStrategy { + PerSubpopulationInstance, // Default + SharedMultiSubpopulation { kind: HydraKind, params: HydraParams }, +} +pub enum HydraKind { HydraKll /* experimental, no error bound */, HydraCms, HydraCountSketch } +pub enum HydraParams { + HydraKll { k: u32, shared_buckets: u32 }, + HydraCms { width: u32, depth: u32, shared_rows: u32, shared_columns: u32 }, + HydraCountSketch { width: u32, depth: u32, shared_rows: u32, shared_columns: u32 }, +} +pub fn hydra_kind_for(a: &SketchAlgorithm) -> Option; // Cms, CountSketch only + +pub enum SamplingKind { Reservoir } pub enum SamplingParams { Reservoir { size: u32 } } +pub enum WaveletKind { Haar } pub enum WaveletParams { Haar { coefficients: u32 } } +pub enum StatModelKind { Parametric } pub enum StatModelParams { Parametric { family: String } } +``` + +### 6.3 Update input and readouts (`post_asap/sketch.rs`, `post_asap/maintained_population.rs`) + +```rust +/// One state update: `item` keys the update for keyed families; `weight` is applied to state. +pub struct SummaryUpdate { + pub item: Option, + pub weight: SummaryInputExpr, + pub weight_domain: WeightDomain, // serde default: UnknownOrSigned +} +impl SummaryUpdate { pub fn column(c: ColumnRef) -> Self; } // item None, UnknownOrSigned +pub enum WeightDomain { + UnknownOrSigned, // Default; never assumed non-negative + NonNegative { proof: NonNegativeWeightProof }, +} +pub enum NonNegativeWeightProof { UnitCount, ResetAwareCounterDerivative } +pub enum SummaryInputExpr { + Constant(f64), Column(ColumnRef), Tuple(Vec), EntityIdentity(EntityIdentity), +} +pub enum EntityIdentity { PromqlLabelSet { excluding: Vec } } + +/// Readout of sketch state, carried by SummaryEstimate. +pub enum SketchStatistic { + FrequencyL2, FrequencyEntropy, + Quantile { q: f64 }, + PointCount { key: ColumnRef, value: Option }, + Cardinality, + TopK { k: usize }, +} + +/// Readout of a maintained population, carried by EvaluatePopulation. +pub enum PopulationStatistic { Quantile { q: f64 }, TopK { k: usize }, Sum, Count, Average } +pub struct MaintainedPopulation { + pub input: PopulationInput, + pub max_k: usize, // largest TopK it supports + pub quantiles: bool, // whether Quantile is supported +} +pub enum PopulationInput { + CurrentSeries(CurrentSeriesInput), // metric, matchers, grouping, without, lookback_ms + Rows { input: Rc, value_column: usize, grouping: GroupKeys }, +} +``` + +A finalized value's accuracy statement is `ResultGuarantee { metric, bound, failure_probability, provenance }` (`post_asap/guarantee.rs`). It is attached to readout and finalized nodes, never to raw state. + +### 6.4 ASAP operators (`crates/types/src/ir/asap.rs`) + +```rust +pub const UNIMPLEMENTED_ASAP_OP: &str = + "this ASAP operator is reserved: schema, accuracy, timing and export are not implemented"; + +pub enum ASAPOp { + SummaryAgg { + child: Rc, + family: FieldDataType, // never Plain + input: SummaryUpdate, + reduction: Reduction, // Reduce(GroupKeys) | PerEntity + grouping: GroupingStrategy, + filter: Option, // serde default None + }, + SummaryEstimate { summary_input: Rc, query: SketchStatistic }, + FinalizeExactAccumulator { child: Rc }, + MaintainPopulation { child: Rc, population: MaintainedPopulation }, + EvaluatePopulation { child: Rc, evaluation: PopulationStatistic }, + // Reserved on this branch; #560 implements SummaryMerge. + SummaryMerge { children: Vec> }, + SummarySubtract { left: Rc, right: Rc }, + SummaryDelete { summary_input: Rc, key: ColumnId }, + SummaryJoin { outer: Rc, inner: Rc, key: ColumnId, family: FieldDataType }, + Extension { child: Rc, name: String }, +} + +impl ASAPOp { + pub fn children(&self) -> Vec<&Rc>; // SummaryAgg includes its filter's subquery nodes + pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Self; + pub fn kind_name(&self) -> &'static str; + /// Merge, Subtract, Delete, Join, Extension on this branch; #560 removes Merge. + pub fn is_unimplemented(&self) -> bool; + /// SummaryAgg/SummaryJoin `family`; #560 adds SummaryMerge (its inputs' state type). + pub fn produced_state(&self) -> Option<&FieldDataType>; + /// #560: merge_disjoint of the children's coverage; fails closed. + pub fn merged_coverage(&self) -> Result; + pub fn output_schema(&self) -> Result; + pub fn output_kind(&self) -> OperatorResultKind; + pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>; +} +``` diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index f87456a9f..f850116bb 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -62,7 +62,7 @@ operator inputs and scalar query-result references use `Rc`. read by expressions. Keep `ScalarExpr::Column(ColumnId)`: the ID selects a field for type checking and the corresponding input value for evaluation, independently of the executor's row/column storage layout. See the -[fields versus column references contract](operator-sharing.md#21-one-schema-model-for-values-and-state). +[fields versus column references contract](asap-primitive-schema.md#21-consideration-1-the-schema-is-the-edge-between-two-nodes). Names are resolved to `ColumnId` before constructing these nodes. Parsing and unresolved `ColumnRef` handling remain frontend concerns; no alternative generic diff --git a/docs/design_docs/proposals/operator-sharing.md b/docs/design_docs/proposals/operator-sharing.md index 1851d1809..c9ab193a4 100644 --- a/docs/design_docs/proposals/operator-sharing.md +++ b/docs/design_docs/proposals/operator-sharing.md @@ -365,155 +365,13 @@ caching or mutation mechanism. ### 2.1 One schema model for values and state -Use one `Schema` for operator outputs before and after optimization. Rename today's -`SummaryFamilyType` to `FieldDataType`: it types every field, and `Plain` is not a summary -family. Rename `Column` to `Field` and `Schema.columns` to `Schema.fields`: the struct -describes a column and holds none of its data. Retain the current `Schema` metadata. -The following is the proposed resolved interface; it is not the current Rust definition. - -```rust -struct Field { - name: String, - dtype: FieldDataType, - nullable: bool, - table: Option, -} - -struct Schema { - fields: Vec, - time_index: Option, - unique_keys: Vec>, - closed: bool, -} - -// Today's `SummaryFamilyType`, renamed; variants and payloads unchanged. -enum FieldDataType { - Plain(DataType), - ExactAggregate(ExactKind, ExactParams), - Sketch(SketchKind, GroupingStrategy), - Sample(SamplingKind, SamplingParams), - Wavelet(WaveletKind, WaveletParams), - StatModel(StatModelKind, StatModelParams), -} - -// Proposed derived output classification, separate from column types. -enum OperatorResultKind { - Relation, - InstantVector, - RangeVector, - State, -} - -impl Operator { - fn output_schema(&self) -> Result; - fn output_kind(&self) -> Result; - fn validate_inputs(&self) -> Result<(), QueryExprError>; -} - -impl OperatorNode { - fn validate_structure(&self) -> Result<(), QueryExprError>; - fn validate_execution_timing(&self) -> Result<(), QueryExprError>; -} - -impl ScalarExpr { - fn scalar_type(&self, input: &Schema) -> Result<(DataType, bool), QueryExprError>; -} -``` - -**Fields versus column references.** These names describe different roles, not -competing representations of the same object: - -| Name | Role | Holds runtime values? | -|---|---|---| -| `Schema` | Ordered `Field` metadata, plus key/time/closedness information | No | -| `Field` | Name, type, nullability and optional qualifier for one output column | No | -| `ColumnRef` | Unresolved logical reference: `Named`, `Qualified`, `SampleValue`, or `Wildcard` | No | -| `ColumnId = usize` | Resolved column position in a particular input/output schema | No | -| Runtime batch | Values conforming to a schema; storage layout is executor-specific | Yes | - -Keep `ColumnRef`, `ColumnId`, and `ScalarExpr::Column(ColumnId)`. Renaming the -metadata struct `Column` to `Field` does not rename column references to field -references. The same position identifies metadata during planning and values -during execution; it is not a stable field identity across projections or joins. -Schema `unique_keys` and `time_index` also use these column positions. - -For example, resolving `t.bytes` to position `1` produces `ColumnId = 1`. -`schema.fields[1]` supplies its type and nullability; evaluating -`ScalarExpr::Column(1)` reads the corresponding value. The native executor -currently reads `row[1]` from `Batch { schema, rows: Vec> }`. A columnar -executor would select array `1` instead. No physical `Column` container is -introduced by the metadata rename, and the old metadata `Column` struct is not -retained as a second type. - -**Relationship to current types.** `Field` is today's pre-ASAP `Column` with `dtype` -widened from `DataType` to `FieldDataType`. `FieldDataType` is today's `SummaryFamilyType` -under a name that also fits its `Plain` case. The proposed common `Schema` replaces -the separate operator-edge roles of pre-ASAP `Schema` and post-ASAP `SummarySchema` / -`SummaryField`; it does not rename `DataType`. A pre-ASAP value column becomes -`Plain(dtype)`. -Frontend validation permits only ordinary value columns, preserving the current -pre-ASAP restriction even though the common schema can also express state. - -| Field | Meaning and requirement | -|---|---| -| `fields` | Ordered named fields. `Plain(DataType)` is a readable value; other variants retain the identity and parameters of summary or exact-accumulator state. | -| `Field.nullable`, `Field.table` | Preserve SQL nullability and qualified column resolution. | -| `time_index` | Identifies the time column when present; it does not by itself distinguish an instant vector from a range vector. | -| `unique_keys` | Proven column combinations identifying rows; an empty list asserts no known key. Recompute these proofs when a rewrite changes identity. | -| `closed` | Whether `fields` completely describes the output. An open PromQL schema must retain unlisted labels through the existing complete-series-identity contract. | - -`OperatorResultKind` is derived from the operation and its inputs and retained as -`OperatorNode.result_kind`. `State` describes an output carrying unfinalized state; its -schema may also contain ordinary grouping keys. `SummaryEstimate`, -`FinalizeExactAccumulator` and other readouts derive the appropriate relation or -vector kind from their operation and input context. Matching numeric columns do -not make those kinds interchangeable. - -**Interface contracts.** `Operator::output_schema` and `output_kind` derive output -metadata from the payload and validated inputs. `validate_inputs` checks local -producer/consumer compatibility, such as vector inputs for `BinaryOp` or the -required state family for a summary readout. Scalar typing checks the input-kind -contract of `PromqlScalarFromVector` and other scalar plan reads. - -| Validation entry | Scope and stage | -|---|---| -| `OperatorNode::validate_structure()` | Walks the reachable operator DAG, including scalar plan references; checks input contracts, scalar typing and agreement between retained and derived output metadata. Valid for logical and physical plans; permits `timing = None`. | -| `OperatorNode::validate_execution_timing()` | Includes structural validation, then requires assigned timing on every executable operator and checks phase dependencies. Used for executable physical candidates. | -| Existing planner assessment and selection (#509) | Establishes guarantees using the existing accuracy models and checks them against request requirements and deployment capabilities. Neither node method re-proves a guarantee or decides request feasibility. | - -The two node methods need only the DAG and its annotations. Request requirements -and deployment models remain inputs to the existing planning/selection workflow, -not implicit globals of `validate_structure`. Passing the timing check alone does -not establish that a physical candidate satisfies the query's accuracy requirement. - -`Scan.schema` declares the source columns; `Values.schema` declares the constructed -row shape. `OperatorNode.schema` is the derived output for any operation. A scan's -predicates cannot change its declared output columns; a Values row must match the -declared arity, types and nullability. These leaf outputs retain the declaration's -column layout and time/identity information, with only justified metadata changes. -The declaration and derived output therefore have distinct roles, and structural -validation rejects disagreement rather than trusting two independent schemas. - -`scalar_type` keeps the existing method name and `(DataType, nullable)` result. -Its `input` is the applicable column scope: the child schema for a projection, -both input schemas for a join predicate, or aggregate outputs for `HAVING`. -Explicit subquery/conversion expressions validate their referenced producer using -the contracts above. Numeric expressions cannot consume state columns as numbers. -A standalone scalar expression is checked with an empty column scope and needs no fabricated -relation output schema. `QueryExprError` retains the existing error-type name; -result-kind, state-family, schema and execution-phase mismatches require -corresponding validation errors. - -For example, a KLL build outputs `State` with a -`Sketch(SketchKind, GroupingStrategy)` column identifying KLL and its parameters. -Its p99 readout outputs an ordinary `Plain(Float64)` column in the appropriate -relation/vector schema. A numeric predicate can use that readout, but not the KLL -state. Exact accumulator state similarly requires `FinalizeExactAccumulator`. -An ordinary operator may pass state through only where its input/output contract -permits it. A bare-column projection can preserve the field's `FieldDataType` -directly during `output_schema` derivation; `scalar_type` applies when that column -is used as a scalar value and rejects state. Copying a state column does not turn -it into a readable scalar. +Every operator output, before and after optimization, uses one `Schema` whose +`Field`s are typed by `FieldDataType`: `Plain(DataType)` for a readable value, or +the family, algorithm and parameters of summary or exact-accumulator state. +`OperatorResultKind` marks state outputs, and state becomes a value only through +an explicit readout. The schema model, `ColumnRef` versus `ColumnId`, the +validation entry points and the readout boundary are specified in +[Schema and physical data for ASAP primitives](asap-primitive-schema.md). ### 2.2 Preserve existing accuracy semantics diff --git a/docs/design_docs/proposals/univmon-frequency-summary.md b/docs/design_docs/proposals/univmon-frequency-summary.md index 75cfd080b..a2d9f96c1 100644 --- a/docs/design_docs/proposals/univmon-frequency-summary.md +++ b/docs/design_docs/proposals/univmon-frequency-summary.md @@ -35,7 +35,8 @@ cardinality alternatives, and exact count remains the cheaper first count candidate. All four readouts have the same unit-weight update, input sub-DAG, grouping, -window, parameter identity and state schema. Existing post-ASAP structural +window, parameter identity and state schema +([ASAP primitive schema](asap-primitive-schema.md)). Existing post-ASAP structural sharing can therefore intern their state producer while preserving distinct readout nodes. Sharing is only legal within the same execution/data scope. Precompute placement, SummaryCatalog installation, retention, and runtime diff --git a/docs/develop_docs/asap-aware-mapping-contracts.md b/docs/develop_docs/asap-aware-mapping-contracts.md index 447cb3823..011758143 100644 --- a/docs/develop_docs/asap-aware-mapping-contracts.md +++ b/docs/develop_docs/asap-aware-mapping-contracts.md @@ -325,24 +325,12 @@ backend inspection. Automatic selection skips those unproven ratios. Use ### Family, category, algorithm, and parameters -Sketches separate their query category from the concrete algorithm and its parameters: +A summary's identity has four levels: family (`FieldDataType` variant), sketch +category (`SketchCategory`), algorithm (`SketchAlgorithm`), and the validated +committed choice (`SketchKind`). The levels and their validation are specified in +[Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#22-consideration-2-a-field-can-have-an-asap-primitive-type). -| Level | Type | Example | -| --- | --- | --- | -| **family** | `SummaryFamilyType` | `Sketch`, `Sample`, `Wavelet`, `StatModel`, `ExactAggregate` | -| **category** | `SketchCategory` | `Quantile`, `Cardinality`, `Frequency`, `TopK` | -| **algorithm** | `SketchAlgorithm` | `Kll` / `DDSketch` (both quantile); `Hll` (HyperLogLog) / `Theta` / `Kmv` (K-Minimum Values), all cardinality | -| **committed choice** | `SketchKind` | one validated category + algorithm + parameter combination | - -A `SketchKind` is a validated committed choice. Its public constructor, -`SketchKind::new(algorithm, params)`, verifies that the parameter variant belongs -to the selected algorithm and classifies the pair into its category. The public -`.category()`, `.algorithm()`, and `.params()` accessors expose the committed -values without permitting an invalid combination. - -Where this matters in practice: `CostModel::rank_candidates`, `CostModel::size_params`, and `SketchAlgorithmStrategy::replacements` operate at the **algorithm** level. `summary_candidates(intent)` returns a list of `SketchAlgorithm`s (`[Kll, DDSketch]` for a `Quantile` intent), never a bare `SketchKind` with nothing chosen underneath it. `SketchKind` appears after an algorithm has been selected and sized—on `Realization::Sketch(SketchKind)` and `SummaryFamilyType::Sketch(SketchKind)`. - -`Sample`, `Wavelet`, and `StatModel` each use a flat `(Kind, Params)` pair. `Sketch` needs the additional algorithm level because multiple algorithms can serve the same purpose—for example, KLL and DDSketch both answer quantile queries. +Where this matters in practice: `CostModel::rank_candidates`, `CostModel::size_params`, and `SketchAlgorithmStrategy::replacements` operate at the **algorithm** level. `summary_candidates(intent)` returns a list of `SketchAlgorithm`s (`[Kll, DDSketch]` for a `Quantile` intent), never a bare `SketchKind` with nothing chosen underneath it. `SketchKind` appears after an algorithm has been selected and sized—on `Realization::Sketch(SketchKind)` and `FieldDataType::Sketch(SketchKind, GroupingStrategy)`. --- @@ -378,7 +366,7 @@ The crate provides no default `Matcher` implementation because the answer depend Concretely, `explanation.rs` reports three candidate kinds from each `TargetSubDAGCandidates`: -- `ExplanationKind::SketchApproximation` — the set contains a `Replacement::Summary` that realizes `SummaryFamilyType::Sketch(..)`, not just an exact/pass-through candidate. +- `ExplanationKind::SketchApproximation` — the set contains a `Replacement::Summary` that realizes `FieldDataType::Sketch(..)`, not just an exact/pass-through candidate. - `ExplanationKind::CommonSubexpressionReuse` — `consumer_count >= 2` and the set contains `SharedSubDAGStrategy`'s "build once and share" candidate (the `Replacement::Rewrite` whose `Rc` is the set's `target`). - `ExplanationKind::ExactComposition` — the candidate set contains an exact operation diff --git a/docs/develop_docs/pre-asap-ir.md b/docs/develop_docs/pre-asap-ir.md index abf5dc50b..2492e28d7 100644 --- a/docs/develop_docs/pre-asap-ir.md +++ b/docs/develop_docs/pre-asap-ir.md @@ -17,26 +17,10 @@ The pre-ASAP IR is defined using the `QueryExpr` enum. We discuss some of import ## Fields and column references -`Schema` owns `Field` metadata: name, type, nullability, and an optional table -qualifier. A `Field` contains no runtime values. The former schema `Column` -struct served this same metadata role; it was renamed to `Field`, not retained -as a second data container. - -`ColumnRef` is an unresolved logical reference (`Named`, `Qualified`, -`SampleValue`, or `Wildcard`). Resolution binds a reference to `ColumnId`, a -`usize` position within a particular schema. `QueryExpr::Column(ColumnId)` -reads that column; the same position indexes `Schema::fields` for type checking -and a runtime row for its value. Group keys, unique keys, and `time_index` also -use these column positions. They are not stable identities across projections -or joins, so the positional reference remains `ColumnId`, not `FieldId`. - -The native runtime names shared ownership `SchemaRef = Arc` and stores -`Batch { schema: SchemaRef, rows: Vec> }`. `Schema` is the same metadata -model during planning and execution; the `Ref` suffix only distinguishes ownership. -It has no physical `Column`/array container. A column reference expresses what -to read independently of whether an executor stores its data as rows or arrays. -For example, resolving `t.bytes` to `ColumnId = 1` obtains its type from -`schema.fields[1]`; native execution reads `row[1]`. +`Schema` holds `Field` metadata (name, type, nullability, qualifier) and no +values; an unresolved `ColumnRef` resolves to a positional `ColumnId` within one +schema. The design, including how the same position selects a runtime value, is +in [Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#21-consideration-1-the-schema-is-the-edge-between-two-nodes). ## Node index From 1f93e4da90a29ec079e4b13f23d744854836900a Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 3 Oct 2026 20:52:04 +0000 Subject: [PATCH 02/59] docs: restore the hand-written sections before the examples MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Keep the author's sections 1–4 as written; the operator examples (§5) and key code interfaces (§6) stay. Co-Authored-By: Claude Opus 5.5 --- .../proposals/asap-primitive-schema.md | 70 ++----------------- 1 file changed, 4 insertions(+), 66 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 8840a892d..b54fe62c4 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -1,7 +1,5 @@ # Schema and Physical Data for ASAP Primitives -> Status: design. `SummaryCoverage` and its rules are implemented in [#567](https://github.com/ProjectASAP/ASAPPlanner/pull/567); `SummaryMerge` in [#560](https://github.com/ProjectASAP/ASAPPlanner/pull/560) (both open). Audience: designers and architects. - This document is the single source of truth for the schema, and column design for ASAP Primitives. This is used in the logical stage (LogicalASAPDAG), and physical stage (PhysicalASAPDAG). ## 1. Goal, problem, and requirements @@ -9,7 +7,7 @@ This document is the single source of truth for the schema, and column design fo Unlike existing Database engines, which work on raw data or explicitly defined materialized tables with schema and column names provided by the users, ASAPPlanner is designed for querying and execution over the mix of raw data and ASAP Primitives. ASAP primitives are usually compact summaries over raw data. Therefore, it introduces new requirement when we design the schema and node definitions for LogicalASAPDAG and PhysicalASAPDAG. Assuming we have the Logical DAG defined for a canonicalized representation for a batch of queries. [TODO: add links for this here. ] -The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG [TODO: link PR 511's doc here], but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: `SummaryAgg` (creates and updates state; there is no separate SummaryCreation or SummaryUpdate operator, and `SummaryUpdate` is `SummaryAgg`'s update-expression parameter), `SummaryEstimate`, `FinalizeExactAccumulator`, `MaintainPopulation` and `EvaluatePopulation` (implemented); `SummaryMerge` (reserved here, enabled by #560); and `SummarySubtract`, `SummaryDelete`, `SummaryJoin` and `Extension` (reserved). See §5. +The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG [TODO: link PR 511's doc here], but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives: SummaryCreation?, SummaryUpdate, SummaryMerge, SummaryDelete, SummarySubtraction, SummaryEstimate [TODO: check what is the complete list or discuss with others about the list]. Each of the Summary operators also require the ASAP primitive information above to inter-operate correctly, preserving semantic correctness. Basically, the following information should be represented to preserve the equivalent query semantics when we introduce ASAP Primitives to logical query representation, and following physical one. @@ -35,78 +33,18 @@ Schema represents the **metadata** of information flow along an **edge** between Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. Based on our requirement, each field should contain the following information. -1. **What type of the ASAP Primitive is** A state column can be a raw data type (e.g., numerical number, string). It can also be a summary type (§6.1, §6.2), e.g., the summary **family** is sketch (`FieldDataType::Sketch`), its **category** is quantile (`SketchCategory::Quantile`), its **algorithm** is KLL (`SketchAlgorithm::Kll`), and its **parameters** are `SketchParams::Kll { k: 200 }`. A sketch field also records its `GroupingStrategy` (one instance per group, or one shared Hydra structure). Non-sketch families (`ExactAggregate`, `Sample`, `Wavelet`, `StatModel`) have a kind and parameters but no category. -2. **What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents** This information is being mapped based on the primitive type: it is not stored in the field. `SummaryEstimate::validate_inputs` accepts a readout only when the `SketchStatistic` matches the state's `SketchCategory` (Quantile→`Quantile`; Cardinality→`Cardinality`/`Universal`; PointCount→`Frequency`/`Universal`; FrequencyL2/FrequencyEntropy→`Universal`; TopK→`TopK`/`Universal`). Exact accumulators are read by `FinalizeExactAccumulator` instead. `Sample`, `Wavelet` and `StatModel` state has no readout operator yet. +1. **What type of the ASAP Primitive is** A state column can be a raw data type (e.g., numerical number, string). It can also be a [summary type](TODO: add link), e.g., the summary family is sketch, and the sketch type is quantile KLL sketch algorithm, and KLL sketch has K as parameter as the schema. (TODO: confirm the terminology with corresponding code/doc) It has a family, an algorithm and parameters. +2. **What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents** This information is being mapped based on the primitive type. -Whether an edge carries **state or values** is recorded on the producing node as `OperatorResultKind`, not inferred from field types. Usually they agree: a `SummaryAgg` output has exactly one non-plain field and kind `State`. They can differ, though. `MaintainPopulation` keeps an all-plain schema but its kind is `State`, and a `Project` that passes a state column through keeps kind `State`. Consumers check `result_kind` (`validate_inputs` requires `State` for every readout). ## 4. Proposed Node field design A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. -In code this field is `OperatorNode::coverage`. It records *which observations* a summary state covers: a source plus a union of joint (time × population) regions. It sits beside the schema, not inside a field, because two states with identical schemas can cover different data, and only disjoint coverage can be merged once-per-observation. - -Based on the above the proposed OperatorNode interface is as below (`crates/types/src/ir/node.rs`, `ir/summary_coverage.rs`): +Based on the above the proposed OperatorNode interface is as below: ```rust -pub enum OperatorResultKind { Relation, InstantVector, RangeVector, /** unfinalized summary/accumulator state */ State } - -pub enum Operator { NonASAP(NonASAPOp), ASAP(ASAPOp) } - -pub struct OperatorNode { - pub operator: Operator, - pub result_kind: OperatorResultKind, // derived from operator + children - pub schema: Schema, // derived; may override names/qualifiers only - pub guarantee: Option, // None until accuracy assessment; None != exact - pub timing: Option, // IngestionTime | QueryTime; None until assigned - #[serde(default)] - pub coverage: Option, // observations a State output summarizes -} - -impl OperatorNode { - pub fn new(op: Operator) -> Result; // derives schema + kind; coverage = None - pub fn with_schema(op: Operator, schema: Schema) -> Self; // caller-supplied names - pub fn new_shared(op: Operator) -> Result, SchemaDerivationError>; - pub fn with_guarantee(self, g: Option) -> Self; - pub fn with_timing(self, t: Option) -> Self; - /// Validates the coverage, then rejects a non-State node (CoverageError::NotState). - pub fn with_coverage(self, c: SummaryCoverage) -> Result; - /// This branch: true only for SummaryAgg. #560: SummaryAgg | SummaryMerge. - pub fn requires_coverage(&self) -> bool; - /// #560: (update expression, reduction) of a SummaryAgg, or shared by a SummaryMerge's inputs. - pub fn summary_update(&self) -> Option<(&SummaryUpdate, &Reduction)>; - /// Rebuilds with new inputs; re-derives schema; clears guarantee, timing and coverage. - pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Result; - /// Per node: coverage well-formed (and present if required), validate_inputs, - /// result_kind and schema structure agree with derivation. - pub fn validate_structure(self: &Rc) -> Result<(), SchemaDerivationError>; - pub fn validate_execution_timing(self: &Rc) -> Result<(), SchemaDerivationError>; - // also: asap(), non_asap(), is_asap(), children(), contains_asap(), reachable() -} - -pub struct SummaryCoverage { - pub source: Source, // Source::Table { table_ref } | Source::TimeSeries { metric } - pub regions: Vec, // union of joint regions, never a Cartesian product -} -pub struct CoverageRegion { - pub time_ms: Option>, // half-open, on the source's time column; None = unrestricted - pub population: BTreeMap, // conjunction of equality predicates; empty = unrestricted -} -impl SummaryCoverage { - /// Intervals non-empty, dimension names non-empty, regions pairwise provably disjoint. - pub fn validate(&self) -> Result<(), CoverageError>; - /// Union of provably disjoint inputs from one source; coalesces adjacent - /// intervals with identical populations and keeps gaps. - pub fn merge_disjoint(inputs: &[Self]) -> Result; -} -pub enum CoverageError { - InvalidInterval, InvalidPopulation, SourceMismatch, PossibleOverlap, EmptyMerge, NotState, Missing, - UnknownInput, // #560: a merge input has no coverage - MergeOutputMismatch, // #560: retained merge coverage differs from the input union -} ``` -Two regions are provably disjoint only if their time ranges do not intersect or they bind the same population dimension to different values. A region without time bounds overlaps any region it is not population-disjoint from. Coverage is caller-established: `validate` checks that it is well-formed, not that it matches the child's predicates. Any rewrite through `map_children` drops it, so the rewriter must declare it again. `#537` adds export and CSE of the logical DAG (`ir/export.rs`, `ir/cse.rs`); no interface in this document depends on it. - ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators Given that these information requirements are introduced by summary operators to work correctly semantically, we show the examples of how the defined OperatorNode, schema, and physical data information work with each kind of summary operators. From 085cae05094e60a3c76c14648e792e3ec8b1b738 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sun, 4 Oct 2026 00:37:40 +0000 Subject: [PATCH 03/59] docs: a top-k readout returns selected rows MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Follows #579: the top-k sketch readout derives one row per selected item instead of a packed Utf8 column, matching the exact Sort → Limit shape. Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index b54fe62c4..fbcfd035f 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -80,7 +80,7 @@ SummaryEstimate(query = SketchStatistic::Quantile { q: 0.99 }) ──Relation(job Utf8, quantile Float64)──▶ (planner may rename to p99) ``` -- Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`), and `topk` Utf8. Keys and metadata pass through. +- Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`). Keys and metadata pass through. A top-k readout is the exception: it returns the selected rows, one per ranked item, with the partition keys, the item identity columns, and a `value` Float64 score (#579). This is the same row shape as an exact Sort → Limit top-k, so the plans for one query share a root schema. - Result kind: the value kind of the source the state was built from (`Relation` here). - Checks: input is `State` with exactly one non-plain field, that field is `Sketch`, and its category accepts the statistic (§3). For example, `Cardinality` on KLL is rejected. - Coverage: **absent**. The output is a value, and `with_coverage` returns `NotState`. From a0db03327c0ccb24a468d90c79b04d63ba993f0a Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Tue, 6 Oct 2026 20:43:49 +0000 Subject: [PATCH 04/59] docs: explain why summary coverage is a node field, not part of the schema Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/proposals/asap-primitive-schema.md | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index fbcfd035f..52603dc79 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -41,6 +41,15 @@ Based on our requirement, each field should contain the following information. A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. +**Why this is a node field, not part of the schema.** Two summary states worth merging always cover different data. `SummaryMerge` requires all inputs to have the same schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2: + +| | State A | State B | Equal? | +|---|---|---|---| +| schema | `(job: Utf8, state: KLL{k=200})` | `(job: Utf8, state: KLL{k=200})` | yes, so the merge is allowed | +| what it summarizes | time `[0,1)` | time `[1,2)` | no, which is why merging them is useful | + +If what a state summarizes were part of the schema, these two schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. So the schema says *what kind of state* this is, and the node field says *which data it was built from*. + Based on the above the proposed OperatorNode interface is as below: ```rust ``` From 7b0a991411dc9f796b64ced39c0e22b84b77b28e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 7 Oct 2026 14:21:07 +0000 Subject: [PATCH 05/59] docs: summary coverage records rows and columns SummaryCoverage records which columns a state summarizes (input, group_by) as well as which rows. Update the SummaryAgg and SummaryMerge examples to #560: merges compare coverage columns, carry no coverage until #646, and merged_coverage is gone. Co-Authored-By: Claude Opus 5.5 --- .../proposals/asap-primitive-schema.md | 51 ++++++++++++++++--- 1 file changed, 43 insertions(+), 8 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 52603dc79..185797a82 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -50,8 +50,44 @@ A node in the physical data will represent the data or summary instance, so a no If what a state summarizes were part of the schema, these two schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. So the schema says *what kind of state* this is, and the node field says *which data it was built from*. +"Which data" has two parts, and the schema records neither: + +- **Rows**: which observations went in: the source, and joint time × population regions. In SQL terms this is `FROM` and `WHERE`. +- **Columns**: which column of each row is fed into the state, and how rows are grouped. In SQL terms this is the argument of the aggregate and `GROUP BY`. A KLL over `latency` by `job` and a KLL over `size` by `job` both have schema `(job: Utf8, state: KLL{k=200})`. + +Two states can merge only when their columns are **identical** and their rows are **disjoint**. Columns that differ would mix latency and size in one state; rows that overlap would count observations twice. + +For `SELECT job, quantile(0.99, latency) FROM t WHERE region = 'us' AND ts in [0, 1min) GROUP BY job`: + +| Coverage field | Records | Value | +|---|---|---| +| `source` | rows | table `t` | +| `regions` | rows | `{ time_ms: 0..60_000, population: { region: "us" } }` | +| `input` | columns | `SummaryUpdate::column(Named("latency"))` | +| `group_by` | columns | `by[job]` | + Based on the above the proposed OperatorNode interface is as below: ```rust +pub struct OperatorNode { + pub operator: Operator, + pub result_kind: OperatorResultKind, + pub schema: Schema, + pub guarantee: Option, + pub timing: Option, + pub coverage: Option, // which data a state summarizes; required on SummaryAgg +} + +pub struct SummaryCoverage { + pub source: Source, // rows: the scanned table or series + pub regions: Vec, // rows: union of joint regions + pub input: SummaryUpdate, // columns: equals SummaryAgg.input + pub group_by: Reduction, // columns: equals SummaryAgg.reduction (child-schema column ids) +} + +pub struct CoverageRegion { + pub time_ms: Option>, // absolute, half-open; None = no time restriction + pub population: BTreeMap, // conjunction of `field = 'text'`; empty = unrestricted +} ``` ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators @@ -71,12 +107,13 @@ SummaryAgg(family = Sketch(KLL{k=200}, PerSubpopulationInstance), input = SummaryUpdate::column(Named("latency")), reduction = by[job], grouping = PerSubpopulationInstance, filter = None) ──State(job Utf8, state Sketch(KLL{k=200}, PerSubpopulationInstance))──▶ - coverage = { source: Table "t", regions: [{ time_ms: 0..60_000, population: {} }] } + coverage = { source: Table "t", regions: [{ time_ms: 0..60_000, population: {} }], + input: column(Named("latency")), group_by: by[job] } ``` - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. - Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **required** and **declared**. `OperatorNode::new` leaves it `None`, `validate_structure` fails with `CoverageError::Missing`, and the planner attaches it with `with_coverage`. +- Coverage: **required** and **declared**. `OperatorNode::new` leaves it `None`, `validate_structure` fails with `CoverageError::Missing`, and the planner attaches it with `with_coverage`. The declared `input`/`group_by` must equal the node's own `input`/`reduction`, or `with_coverage` (and `validate_structure`) fails with `ColumnMismatch`. #646 derives it instead. - Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. ### 5.2 `SummaryEstimate`: sketch state → value @@ -156,12 +193,12 @@ Scenario: combine two one-minute KLL panes over `Scan(t: value Float64)` into a SummaryAgg(KLL k=200, column(SampleValue), by[]) ──State(state Sketch(KLL{k=200}))── coverage {t, [0..60_000)} ─┐ SummaryAgg(KLL k=200, column(SampleValue), by[]) ──State(state Sketch(KLL{k=200}))── coverage {t, [60_000..120_000)} ─┴▶ SummaryMerge - ──State(state Sketch(KLL{k=200}))──▶ coverage = { t, [0..120_000) } (derived) + ──State(state Sketch(KLL{k=200}))──▶ coverage: none until #646 ``` - Output schema: `children[0].schema`. -- Checks: at least one input; exactly one state column; every input is `State` with an identical schema (so family, params, grouping strategy and key positions match); every input has the same `summary_update()` (update expression and reduction); and `merged_coverage()` succeeds. Merging k=200 with k=300 fails, and so does merging raw rows. -- Coverage: **required** and **derived**. `OperatorNode::new` sets it to `SummaryCoverage::merge_disjoint` of the input coverages. An input without coverage gives `UnknownInput`, and overlapping inputs give `PossibleOverlap`. `validate_structure` rejects a retained coverage that differs from the derived one (`MergeOutputMismatch`). Gapped inputs stay as two regions. +- Checks: at least one input; exactly one state column; every input is `State` with an identical schema (so family, params, grouping strategy and key positions match); every input carries coverage (else `UnknownInput`), and all inputs have the same coverage columns, `input` and `group_by` (else `ColumnMismatch`). Merging k=200 with k=300 fails, and so do merging raw rows and merging a KLL over `latency` with one over `size`. +- Coverage: **absent** in #560. #560 does not derive coverage or check that the input rows are disjoint, so a merge whose input is another merge is rejected with `UnknownInput`. #646 derives coverage for every summary node with one `SummaryCoverage::derive`: for a merge, the disjoint union of the input rows (overlapping inputs give `PossibleOverlap`; gapped inputs stay as two regions) with the shared columns. - Boundary: state in, state out. No value is produced until a readout. ### 5.7 Reserved operators (not implemented) @@ -184,7 +221,7 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | `FinalizeExactAccumulator` | `State` (`ExactAggregate`) | source's value kind | no | absent | implemented | | `MaintainPopulation` | `Relation` (table) / `InstantVector` (series) | `State` | yes (by kind; fields plain) | optional, not required | implemented | | `EvaluatePopulation` | `State` from `MaintainPopulation` | source's value kind | no | absent | implemented | -| `SummaryMerge` | `State` × N | `State` | yes | required, derived | reserved; enabled by #560 | +| `SummaryMerge` | `State` × N | `State` | yes | absent in #560; derived in #646 | reserved; enabled by #560 | | `SummarySubtract` | `State` × 2 | `State` | yes | — | reserved | | `SummaryDelete` | `State` | `State` | yes | — | reserved | | `SummaryJoin` | `State` × 2 | `State` | yes | — | reserved | @@ -372,8 +409,6 @@ impl ASAPOp { pub fn is_unimplemented(&self) -> bool; /// SummaryAgg/SummaryJoin `family`; #560 adds SummaryMerge (its inputs' state type). pub fn produced_state(&self) -> Option<&FieldDataType>; - /// #560: merge_disjoint of the children's coverage; fails closed. - pub fn merged_coverage(&self) -> Result; pub fn output_schema(&self) -> Result; pub fn output_kind(&self) -> OperatorResultKind; pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>; From 63acde8f824cf52f10afe0f2545621c8b516019e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 7 Oct 2026 22:34:20 +0000 Subject: [PATCH 06/59] docs: summary coverage as definition + selection, based on Goldstein-Larson view matching Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 236 ++++++++++++++---- 1 file changed, 186 insertions(+), 50 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 185797a82..ee935dae1 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -41,32 +41,117 @@ Based on our requirement, each field should contain the following information. A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. -**Why this is a node field, not part of the schema.** Two summary states worth merging always cover different data. `SummaryMerge` requires all inputs to have the same schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2: +**Why coverage is not part of the schema.** Two summary states worth merging always cover different data. `SummaryMerge` requires all inputs to have the same schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2: | | State A | State B | Equal? | |---|---|---|---| | schema | `(job: Utf8, state: KLL{k=200})` | `(job: Utf8, state: KLL{k=200})` | yes, so the merge is allowed | | what it summarizes | time `[0,1)` | time `[1,2)` | no, which is why merging them is useful | -If what a state summarizes were part of the schema, these two schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. So the schema says *what kind of state* this is, and the node field says *which data it was built from*. +If what a state summarizes were part of the schema, these two schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. So the schema says *what kind of state* this is, and coverage says *which data it was built from*. -"Which data" has two parts, and the schema records neither: +A summary state summarizes the result of a whole computation, not a few columns of a raw table. A KLL over `rate(requests_total[5m])` summarizes rate outputs, and a KLL over a join summarizes join rows. So coverage describes the state by the sub-DAG below it, split into the part that says *what is computed* and the part that says *which of its rows were taken*. -- **Rows**: which observations went in: the source, and joint time × population regions. In SQL terms this is `FROM` and `WHERE`. -- **Columns**: which column of each row is fed into the state, and how rows are grouped. In SQL terms this is the argument of the aggregate and `GROUP BY`. A KLL over `latency` by `job` and a KLL over `size` by `job` both have schema `(job: Utf8, state: KLL{k=200})`. +### 4.1 Design basis: view matching (Goldstein & Larson) -Two states can merge only when their columns are **identical** and their rows are **disjoint**. Columns that differ would mix latency and size in one state; rows that overlap would count observations twice. +The design follows the view matching algorithm of Goldstein and Larson, which decides when a query can be answered from a materialized select-project-join-group-by (SPJG) view: -For `SELECT job, quantile(0.99, latency) FROM t WHERE region = 'us' AND ts in [0, 1min) GROUP BY job`: +> J. Goldstein and P.-Å. Larson. *Optimizing Queries Using Materialized Views: A Practical, Scalable Solution.* SIGMOD 2001. -| Coverage field | Records | Value | +The algorithm splits a view's `WHERE` into column equivalence classes, a **range** per column and **residual** predicates. A view can answer a query when the residuals match, the query's ranges lie inside the view's (§3.1.2), the columns needed by compensating predicates are in the view output (§3.3, requirement 2), and the query's `GROUP BY` is a subset of the view's, so the query's groups are further aggregations of the view's groups (§3.3, requirement 3). The SPJ part is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`); it rejects `Aggregate` and `Join` input plans. + +A summary state is an aggregation view whose aggregate is a summary family. The mapping is: + +| Goldstein & Larson | Summary coverage | +|---|---| +| SPJ part: tables, joins, residual predicates | the computation `C` below the `SummaryAgg` (§4.2), part of `definition` | +| aggregate function and its argument | `family` and `input` (`SummaryUpdate`) of the `SummaryAgg`, part of `definition` | +| `GROUP BY` | the `SummaryAgg` reduction `G`, part of `definition` | +| ranges per column | `selection` (§4.3) | +| compensating predicate on view output | slice on a column of `G` only (§4.4) | +| query `GROUP BY` ⊆ view `GROUP BY` | rollup (§4.4) | + +What this design adds beyond the paper: + +- **Unions of states.** The paper considers single-view substitutes and notes that requirement 1 "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. +- **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.4). +- **Value sets and hash partitions** next to ranges, and **evaluation-relative time** (§4.3). + +### 4.2 Coverage = definition + selection + +A state built by `SummaryAgg` means + +```text +state_g = family( input( σ( C ) ) ) for each group value g of G, restricted to G = g +``` + +- `C` is the child sub-DAG with the selection removed. Its output rows are the contributions. +- `σ` is the selection: which output rows of `C` went into the state. + +Coverage stores exactly these two things: + +- **`definition`**: the `SummaryAgg` node itself, with the selection removed from its child sub-DAG. It carries `C`, `input`, `family` and `G`. It is what the state *means*. +- **`selection`**: a union of boxes over the output columns of `C`. It is *which rows* the state took. + +If two states have the same `definition`, their contributions come from the same rows of the same computation, whatever `C` contains (join, union, `rate`, dedup). Disjoint selections then cannot share a row, so no observation is counted twice. No per-operator occurrence rule is needed. + +Examples of what ends up where: + +| Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | |---|---|---| -| `source` | rows | table `t` | -| `regions` | rows | `{ time_ms: 0..60_000, population: { region: "us" } }` | -| `input` | columns | `SummaryUpdate::column(Named("latency"))` | -| `group_by` | columns | `by[job]` | +| `Filter(region = 'us', Scan t)` | `Scan t` | `region ∈ {us}` | +| `Filter(latency < 100, Scan t)` | `Scan t` | `latency ∈ (−∞, 100)` | +| `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | time `(−3m, −2m]` relative to evaluation | +| `Filter(rate > 0, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))`, `rate > 0` as residual | — | +| `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | + +In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `rate` and changes the rate values, so it is not a selection of output rows. The 5-minute read window is a source dependency, not coverage (§4.6). + +### 4.3 Deriving the selection + +Coverage is derived from the node, never declared. Walking down from the `SummaryAgg` (its own `filter` included), a predicate conjunct goes into `selection` when both hold: + +1. **It can be lifted to the `SummaryAgg`.** Lifting is the inverse of DataFusion's `PushDownFilter` (`datafusion-optimizer`, `push_down_filter.rs`): a predicate passes `Filter`, `TimeRange`/`TimeShift` and a direct-column `Project` (renaming the column); passes an `Aggregate` only when every column it uses is a group column; and passes a window function or a per-series temporal function such as `rate` only when every column it uses is a partition column (a series label). Anything else stops it. +2. **It is one of the box constraints.** Per column, one of: + - **value set**: `In` or `NotIn` a set of literals, from `=`, `!=`, `IN`, `NOT IN` and `OR` of equalities (as DataFusion's `LiteralGuarantee` extracts them); + - **interval**: lower and upper `std::ops::Bound` (`Included`, `Excluded` or `Unbounded`) from comparisons (as DataFusion's `Interval`); + - **hash partition**: `hash(columns) mod n = k`. + +A conjunct that fails either rule stays in `definition` as a residual, as in Goldstein & Larson. Column equalities (`a = b`) are residuals too: there are no column equivalence classes. + +Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns, and a direct alias keeps the identity of the column it renames. + +**Time** is a selection like any other, with two coordinate kinds: + +- A `TimeRange(w)` over a `TimeShift(s)` on the lifted chain gives time **relative to evaluation**: `(Excluded(−(s+w)), Included(−s))`. PromQL ranges are left-open, matching the executor (`series_window.rs`). This is how Stage 2 tumbling panes are built (`window_composition.rs` in #601), so their time is derived rather than declared. +- An interval filter on the timestamp column (the schema's `time_index`) gives **absolute** time, for example `ts >= t0 AND ts < t1` gives `(Included(t0), Excluded(t1))`. + +A relative and an absolute interval are never compared: two states whose times are of different kinds are treated as possibly overlapping. Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. + +### 4.4 Operations + +| Operation | Example | Valid when | +|---|---|---| +| merge (`SummaryMerge`, same `G`) | `[0,1m)` ⊕ `[1m,2m)`; `region='us'` ⊕ `region='eu'` | all `definition`s equal; selections related as the family requires (below) | +| rollup (`SummaryMerge` with `group_by: G'`) | `by[region, job]` → `by[job]` | `G'` ⊆ `G` and the family merges. Groups of one state are disjoint because a row has one value per group column, so no selection check is needed | +| slice | `by[region, job]` state answering `region = 'us' … by[job]` | the restricted columns are all in `G`. A sketch cannot be filtered, so a restriction on any other column is invalid | +| reuse for a query | a stored state answers a query | Goldstein & Larson containment: same `definition`, query selection inside the state's, any compensating restriction is a slice | +| subtract (`SummarySubtract`, reserved) | `[0,10) − [0,5)` | same `definition`; the right selection is contained in the left | + +One `SummaryMerge { children, group_by }` covers both merge and rollup: one child with a coarser `group_by` is a rollup, and `group_by` equal to the children's is a plain merge. Its coverage is the children's `definition` with `G'` and the union of their selections; adjacent intervals are joined, gaps stay as separate boxes. + +How selections must relate is declared by the family, next to whether it merges (`FieldDataType::family_merges`, added in #592): + +- **disjoint** for counting families (KLL, Count-Min, exact `Sum`/`Count`): an overlapping row would be counted twice; +- **overlap allowed** for idempotent families (HLL, exact `Min`/`Max`, distinct sets); +- **contained** for subtraction. + +Two `definition`s are equal when their canonical forms (`canonicalize`) are structurally equal, ignoring planning metadata: `timing`, `guarantee` and `coverage_cache`. `SummaryUpdate.weight_domain` is compared: it is derived from `C` and `input`, so it differs only if a derivation is wrong. A state built at ingestion time and one built at query time can therefore merge. + +States over different sources have different `definition`s and do not merge. To combine tables, put `UNION ALL` with a marker column below one `SummaryAgg`; the marker is then an ordinary column for `selection` or `G`. + +### 4.5 Interface -Based on the above the proposed OperatorNode interface is as below: ```rust pub struct OperatorNode { pub operator: Operator, @@ -74,22 +159,55 @@ pub struct OperatorNode { pub schema: Schema, pub guarantee: Option, pub timing: Option, - pub coverage: Option, // which data a state summarizes; required on SummaryAgg + /// Cache for `coverage()`. Lazily filled, never serialized, ignored by + /// equality. Not a source of truth: coverage is always re-derivable. + coverage_cache: CoverageCache, +} + +impl OperatorNode { + /// `Some` for `SummaryAgg` and `SummaryMerge`. + pub fn coverage(&self) -> Option<&SummaryCoverage>; } pub struct SummaryCoverage { - pub source: Source, // rows: the scanned table or series - pub regions: Vec, // rows: union of joint regions - pub input: SummaryUpdate, // columns: equals SummaryAgg.input - pub group_by: Reduction, // columns: equals SummaryAgg.reduction (child-schema column ids) + /// The `SummaryAgg` (or rolled-up equivalent) with the selection removed. + pub definition: Rc, + /// Union of boxes over the output rows of the definition's computation. + pub selection: Vec, +} + +pub struct SelectionBox { + pub columns: BTreeMap, // (table, name) lineage + pub time: Option, // None = unrestricted +} + +pub enum Constraint { + In(BTreeSet), + NotIn(BTreeSet), + Interval { lower: Bound, upper: Bound }, + HashPartition { columns: Vec, of: u32, index: u32 }, } -pub struct CoverageRegion { - pub time_ms: Option>, // absolute, half-open; None = no time restriction - pub population: BTreeMap, // conjunction of `field = 'text'`; empty = unrestricted +pub enum CoverageTime { + RelativeToEvaluation { lower: Bound, upper: Bound }, + Absolute { lower: Bound, upper: Bound }, +} + +impl SummaryCoverage { + pub fn derive(node: &OperatorNode) -> Result; } ``` +`OperatorNode::new` still rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. + +### 4.6 What coverage does not contain + +- **Source dependencies**: which source rows must be read to compute the contributions, such as the 5-minute window under `rate`. This is read planning and maintenance (compare `materialized/dependencies.rs` in `datafusion-materialized-views`). +- **Readiness and completeness**: whether a stored instance holds all of its rows. +- **Absolute binding** of relative time, and deployment identity. + +These belong to ASAPQuery-backend. The SDS split matches coverage: `SummaryDefinition` stores the serialized `definition` (Planner provides its serde; Backend owns the format version, definition id and hash), and a `StoredSummary`'s coordinates are the `selection` bound to one evaluation plus the group value. + ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators Given that these information requirements are introduced by summary operators to work correctly semantically, we show the examples of how the defined OperatorNode, schema, and physical data information work with each kind of summary operators. @@ -98,22 +216,25 @@ Notation: an edge is written `──Kind(field Type, …)──▶`. Schemas are ### 5.1 `SummaryAgg`: values → state -Scenario: p99 latency by job, from KLL(k=200), over one minute of table `t`. +Scenario: p99 latency by job, from KLL(k=200), over one minute of table `t`, US rows only. ```text -Scan(t: job Utf8, latency Float64) - ──Relation(job Utf8, latency Float64)──▶ +Scan(t: job Utf8, region Utf8, ts Int64 [time_index], latency Float64) + ──Relation(job Utf8, region Utf8, ts Int64, latency Float64)──▶ +Filter(region = 'us' AND ts >= 0 AND ts < 60_000) + ──Relation(job Utf8, region Utf8, ts Int64, latency Float64)──▶ SummaryAgg(family = Sketch(KLL{k=200}, PerSubpopulationInstance), input = SummaryUpdate::column(Named("latency")), reduction = by[job], grouping = PerSubpopulationInstance, filter = None) ──State(job Utf8, state Sketch(KLL{k=200}, PerSubpopulationInstance))──▶ - coverage = { source: Table "t", regions: [{ time_ms: 0..60_000, population: {} }], - input: column(Named("latency")), group_by: by[job] } + coverage() = { definition: this SummaryAgg over Scan(t) (the Filter removed), + selection: [{ columns: { t.region: In{'us'} }, + time: Absolute [Included(0), Excluded(60_000)) }] } ``` - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. - Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **required** and **declared**. `OperatorNode::new` leaves it `None`, `validate_structure` fails with `CoverageError::Missing`, and the planner attaches it with `with_coverage`. The declared `input`/`group_by` must equal the node's own `input`/`reduction`, or `with_coverage` (and `validate_structure`) fails with `ColumnMismatch`. #646 derives it instead. +- Coverage: **always derived**, never declared (§4.3). Both conjuncts of the `Filter` lift into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. - Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. ### 5.2 `SummaryEstimate`: sketch state → value @@ -129,7 +250,7 @@ SummaryEstimate(query = SketchStatistic::Quantile { q: 0.99 }) - Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`). Keys and metadata pass through. A top-k readout is the exception: it returns the selected rows, one per ranked item, with the partition keys, the item identity columns, and a `value` Float64 score (#579). This is the same row shape as an exact Sort → Limit top-k, so the plans for one query share a root schema. - Result kind: the value kind of the source the state was built from (`Relation` here). - Checks: input is `State` with exactly one non-plain field, that field is `Sketch`, and its category accepts the statistic (§3). For example, `Cardinality` on KLL is rejected. -- Coverage: **absent**. The output is a value, and `with_coverage` returns `NotState`. +- Coverage: **none**. The output is a value; `coverage()` returns `None`. - Boundary: state is consumed and a value is produced; `guarantee` on this node carries the readout's error bound. ### 5.3 `FinalizeExactAccumulator`: exact state → value @@ -140,14 +261,14 @@ Scenario: total bytes by host with an exact Sum accumulator. Scan(t: host Utf8, bytes Float64) ──Relation(host Utf8, bytes Float64)──▶ SummaryAgg(family = ExactAggregate(Sum, Sum), input = column(Named("bytes")), reduction = by[host]) - ──State(host Utf8, state ExactAggregate(Sum, Sum))──▶ coverage: required, declared + ──State(host Utf8, state ExactAggregate(Sum, Sum))──▶ coverage: derived FinalizeExactAccumulator ──Relation(host Utf8, state Float64)──▶ ``` - Output schema: each `ExactAggregate` field keeps its name (`state`) and takes the type and nullability the equivalent `NonASAPOp::Aggregate` would give: Sum/Min/Max follow the input column, Count is Int64, and Rate/IRate/Increase are Float64. If the child is not a `SummaryAgg` directly, Count falls back to Int64 and the others to Float64. `unique_keys`, `closed` and `time_index` are preserved (`schema_rebuilding.rs`). - Checks: the input is `State` and contains an `ExactAggregate` field; a sketch is rejected (`structure_contract.rs`). -- Coverage: **absent** on the output. +- Coverage: **none** on the output. - Boundary: this is the explicit maintenance-to-read boundary for exact state. Exact state is never read through `SummaryEstimate`. ### 5.4 `MaintainPopulation`: values → maintained membership (state) @@ -165,7 +286,7 @@ MaintainPopulation(population = MaintainedPopulation { - Output schema: identical to the child's, all plain. Only `result_kind = State` marks it as maintained state. - Checks: `population.matches_node(child)`. For `Rows`, the child must be the same closed table `Scan`, the value column must be non-null Float64, and grouping must be `by` with in-range keys. For `CurrentSeries`, it must be a `TimeSeries` scan with the same metric, matchers and grouping labels, under an instant `TimeRange` of `lookback_ms` (which may be omitted only for the default 300 s lookback). -- Coverage: **not required**. `with_coverage` accepts it because the output is `State`. +- Coverage: **none**. Maintained membership is not combined by `SummaryMerge`. If maintained populations are later materialized per pane, they derive coverage the same way as `SummaryAgg`. - Boundary: the output is state because it must also track membership changes; downstream operators can only read it through `EvaluatePopulation`. ### 5.5 `EvaluatePopulation`: maintained membership → value @@ -180,25 +301,40 @@ EvaluatePopulation(evaluation = PopulationStatistic::Quantile { q: 0.99 }) - Output schema: the schema of `Aggregate(by grouping, measure)` over the maintained source. Quantile gives `quantile_` Float64, Sum gives `sum` (value type), Count gives `count` Int64 and Average gives `avg` Float64; `unique_keys = [[0]]`, `closed`. `TopK { k }` instead returns the source schema unchanged (the selected rows). - Checks: the child is a `MaintainPopulation` node whose `supports(evaluation)` holds: `quantiles` must be set for `Quantile`, and `k <= max_k` for `TopK`. -- Coverage: **absent**. +- Coverage: **none**. - Boundary: maintained membership is read as a value; the result kind is the source's (`Relation`). -### 5.6 `SummaryMerge` (#560): state × N → state +### 5.6 `SummaryMerge`: state × N → state (merge and rollup) + +On `main`, `SummaryMerge { children }` is **reserved**: `is_unimplemented()` returns true, and `output_schema()` and `validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so `OperatorNode::new` fails. #560 enables it for children with identical schemas. This design adds `group_by`, so one operator does both merge and rollup (§4.4): -On this branch, `SummaryMerge { children }` is **reserved**. `is_unimplemented()` returns true, and `output_schema()` and `validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so `OperatorNode::new` fails. Only `output_kind()` (= `State`), `children`, `map_children` and `kind_name` work. #560 enables it as follows (`summary_merge_structure.rs` in #560). +```rust +SummaryMerge { children: Vec, group_by: Reduction } +``` -Scenario: combine two one-minute KLL panes over `Scan(t: value Float64)` into a two-minute state. +Scenario A, time panes: two one-minute KLL panes of PromQL `quantile_over_time(0.99, m[2m])` merged into the two-minute state. Each pane reads `TimeRange(1m)` over `TimeShift(s)` over the scan, as Stage 2 builds them. ```text -SummaryAgg(KLL k=200, column(SampleValue), by[]) ──State(state Sketch(KLL{k=200}))── coverage {t, [0..60_000)} ─┐ -SummaryAgg(KLL k=200, column(SampleValue), by[]) ──State(state Sketch(KLL{k=200}))── coverage {t, [60_000..120_000)} ─┴▶ -SummaryMerge - ──State(state Sketch(KLL{k=200}))──▶ coverage: none until #646 +pane 0 = SummaryAgg(KLL k=200, column(SampleValue), by[]) over TimeRange(1m, TimeShift(0, Scan m)) + coverage() = { definition: SummaryAgg(...) over Scan m, selection: [{ time: Relative (−1m, 0] }] } +pane 1 = SummaryAgg(KLL k=200, column(SampleValue), by[]) over TimeRange(1m, TimeShift(1m, Scan m)) + coverage() = { definition: same, selection: [{ time: Relative (−2m, −1m] }] } +SummaryMerge(children = [pane 0, pane 1], group_by = by[]) + ──State(state Sketch(KLL{k=200}))──▶ + coverage() = { definition: same, selection: [{ time: Relative (−2m, 0] }] } ``` -- Output schema: `children[0].schema`. -- Checks: at least one input; exactly one state column; every input is `State` with an identical schema (so family, params, grouping strategy and key positions match); every input carries coverage (else `UnknownInput`), and all inputs have the same coverage columns, `input` and `group_by` (else `ColumnMismatch`). Merging k=200 with k=300 fails, and so do merging raw rows and merging a KLL over `latency` with one over `size`. -- Coverage: **absent** in #560. #560 does not derive coverage or check that the input rows are disjoint, so a merge whose input is another merge is rejected with `UnknownInput`. #646 derives coverage for every summary node with one `SummaryCoverage::derive`: for a merge, the disjoint union of the input rows (overlapping inputs give `PossibleOverlap`; gapped inputs stay as two regions) with the shared columns. +Scenario B, populations: `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1) merge into `selection: [{ t.region: In{'us', 'eu'} }]`. + +Scenario C, rollup: one `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Every job's state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. + +- Output schema: the children's schema with the group key fields reduced to `group_by`. +- Checks: + - at least one child, every child is `State` with exactly one state field; + - all children have equal `definition`s (§4.4), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; + - `group_by` ⊆ the children's `G`, and the family merges; + - the children's selections relate as the family requires: disjoint for KLL, so pane 0 with pane 0 is rejected; overlap is allowed for HLL. +- Coverage: **derived**: the shared `definition` with `group_by`, and the union of the children's selections. Adjacent intervals join; gaps stay as separate boxes. Nested merges work because a child merge has coverage like any other summary node. - Boundary: state in, state out. No value is produced until a readout. ### 5.7 Reserved operators (not implemented) @@ -207,7 +343,7 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | Operator | Fields | Intended edge shape | |---|---|---| -| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5) | +| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.4) | | `SummaryDelete` | `summary_input, key: ColumnId` | State → State with the entries for `key` removed | | `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | | `Extension` | `child, name` | deployment-named state operator | @@ -216,13 +352,13 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | Operator | Input kind | Output kind | Output carries state | Coverage on output | Status | |---|---|---|---|---|---| -| `SummaryAgg` | value (not `State`) | `State` | yes (one `family` field) | required, declared | implemented | -| `SummaryEstimate` | `State` (one `Sketch` field) | source's value kind | no | absent | implemented | -| `FinalizeExactAccumulator` | `State` (`ExactAggregate`) | source's value kind | no | absent | implemented | -| `MaintainPopulation` | `Relation` (table) / `InstantVector` (series) | `State` | yes (by kind; fields plain) | optional, not required | implemented | -| `EvaluatePopulation` | `State` from `MaintainPopulation` | source's value kind | no | absent | implemented | -| `SummaryMerge` | `State` × N | `State` | yes | absent in #560; derived in #646 | reserved; enabled by #560 | -| `SummarySubtract` | `State` × 2 | `State` | yes | — | reserved | +| `SummaryAgg` | value (not `State`) | `State` | yes (one `family` field) | derived: itself minus selection, plus selection | implemented | +| `SummaryEstimate` | `State` (one `Sketch` field) | source's value kind | no | none | implemented | +| `FinalizeExactAccumulator` | `State` (`ExactAggregate`) | source's value kind | no | none | implemented | +| `MaintainPopulation` | `Relation` (table) / `InstantVector` (series) | `State` | yes (by kind; fields plain) | none | implemented | +| `EvaluatePopulation` | `State` from `MaintainPopulation` | source's value kind | no | none | implemented | +| `SummaryMerge` | `State` × N | `State` | yes | derived: shared definition with `group_by`, union of selections | reserved; enabled by #560, `group_by` added by this design | +| `SummarySubtract` | `State` × 2 | `State` | yes | derived: left selection minus right (planned) | reserved | | `SummaryDelete` | `State` | `State` | yes | — | reserved | | `SummaryJoin` | `State` × 2 | `State` | yes | — | reserved | | `Extension` | any | `State` | yes | — | reserved | From fc5e9a9b2d2237f71818e78d3f9723ec0f909204 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 7 Oct 2026 22:50:04 +0000 Subject: [PATCH 07/59] docs: the 5.1 example's time index is a Timestamp column Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index ee935dae1..8788e15c1 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -219,10 +219,10 @@ Notation: an edge is written `──Kind(field Type, …)──▶`. Schemas are Scenario: p99 latency by job, from KLL(k=200), over one minute of table `t`, US rows only. ```text -Scan(t: job Utf8, region Utf8, ts Int64 [time_index], latency Float64) - ──Relation(job Utf8, region Utf8, ts Int64, latency Float64)──▶ +Scan(t: job Utf8, region Utf8, ts Timestamp [time_index], latency Float64) + ──Relation(job Utf8, region Utf8, ts Timestamp, latency Float64)──▶ Filter(region = 'us' AND ts >= 0 AND ts < 60_000) - ──Relation(job Utf8, region Utf8, ts Int64, latency Float64)──▶ + ──Relation(job Utf8, region Utf8, ts Timestamp, latency Float64)──▶ SummaryAgg(family = Sketch(KLL{k=200}, PerSubpopulationInstance), input = SummaryUpdate::column(Named("latency")), reduction = by[job], grouping = PerSubpopulationInstance, filter = None) From b1bc9c7e154c9a38b011f6fd160bba6cc9d43bfb Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 7 Oct 2026 23:01:18 +0000 Subject: [PATCH 08/59] docs: absolute time is a timestamp-column interval; match the coverage interface in #646 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 33 ++++++++++--------- 1 file changed, 17 insertions(+), 16 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 8788e15c1..f9ffc0fb7 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -121,12 +121,12 @@ A conjunct that fails either rule stays in `definition` as a residual, as in Gol Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns, and a direct alias keeps the identity of the column it renames. -**Time** is a selection like any other, with two coordinate kinds: +**Time** is a selection like any other: -- A `TimeRange(w)` over a `TimeShift(s)` on the lifted chain gives time **relative to evaluation**: `(Excluded(−(s+w)), Included(−s))`. PromQL ranges are left-open, matching the executor (`series_window.rs`). This is how Stage 2 tumbling panes are built (`window_composition.rs` in #601), so their time is derived rather than declared. -- An interval filter on the timestamp column (the schema's `time_index`) gives **absolute** time, for example `ts >= t0 AND ts < t1` gives `(Included(t0), Excluded(t1))`. +- **Absolute** time needs nothing special: it is an interval on the timestamp column (the schema's `time_index`), for example `ts >= t0 AND ts < t1` gives `(Included(t0), Excluded(t1))` on `ts`. The IR has no timestamp literal yet, so such SQL filters stay residual until it does. +- **Relative** time has its own field, because it is not a column value: a `TimeRange(w)` over a `TimeShift(s)` on the lifted chain gives `(Excluded(−(s+w)), Included(−s))` relative to evaluation. PromQL ranges are left-open, matching the executor (`series_window.rs`). This is how Stage 2 tumbling panes are built (`window_composition.rs` in #601), so their time is derived rather than declared. Time is lifted only from a single range `TimeRange`; an instant `TimeRange` picks the latest sample per series, which is not a selection of rows, so it stays in `definition`. -A relative and an absolute interval are never compared: two states whose times are of different kinds are treated as possibly overlapping. Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. +Relative time and a timestamp-column interval are different dimensions, so they are never compared: two states restricted only by different kinds of time are treated as possibly overlapping. Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. ### 4.4 Operations @@ -160,12 +160,13 @@ pub struct OperatorNode { pub guarantee: Option, pub timing: Option, /// Cache for `coverage()`. Lazily filled, never serialized, ignored by - /// equality. Not a source of truth: coverage is always re-derivable. + /// equality, emptied on clone. Not a source of truth: coverage is always + /// re-derivable. coverage_cache: CoverageCache, } impl OperatorNode { - /// `Some` for `SummaryAgg` and `SummaryMerge`. + /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`. pub fn coverage(&self) -> Option<&SummaryCoverage>; } @@ -177,20 +178,20 @@ pub struct SummaryCoverage { } pub struct SelectionBox { - pub columns: BTreeMap, // (table, name) lineage - pub time: Option, // None = unrestricted + pub columns: BTreeMap, // missing column = unrestricted + pub relative_time: Option<(Bound, Bound)>, // ms from evaluation; None = unrestricted } -pub enum Constraint { - In(BTreeSet), - NotIn(BTreeSet), - Interval { lower: Bound, upper: Bound }, - HashPartition { columns: Vec, of: u32, index: u32 }, +pub struct ColumnIdentity { + pub table: Option, + pub name: String, } -pub enum CoverageTime { - RelativeToEvaluation { lower: Bound, upper: Bound }, - Absolute { lower: Bound, upper: Bound }, +pub enum Constraint { + In(Vec), // ScalarValue has no total order (Float64) + NotIn(Vec), + Interval { lower: Bound, upper: Bound }, + // HashPartition { columns, of, index }: added with its first producer. } impl SummaryCoverage { From 832cef8cf6572d546eeb48d224eedc35411a3064 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 7 Oct 2026 23:07:38 +0000 Subject: [PATCH 09/59] docs: ambiguous column names and literal types in selections Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index f9ffc0fb7..d511625d9 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -119,7 +119,7 @@ Coverage is derived from the node, never declared. Walking down from the `Summar A conjunct that fails either rule stays in `definition` as a residual, as in Goldstein & Larson. Column equalities (`a = b`) are residuals too: there are no column equivalence classes. -Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns, and a direct alias keeps the identity of the column it renames. +Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns, and a direct alias keeps the identity of the column it renames. A column whose `(table, name)` is not unique in the output (two items aliased `k`) cannot be named, so its conjuncts stay residual. Value sets compare literals by type: `1` and `1.0` are never proven different. **Time** is a selection like any other: From b1d5817b821023a6f739d090ea3c75269cea2e7b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:43:55 +0000 Subject: [PATCH 10/59] =?UTF-8?q?docs:=20fill=20=C2=A71=20links=20and=20op?= =?UTF-8?q?erator=20list,=20=C2=A72=20terminology,=20=C2=A73=20summary=20t?= =?UTF-8?q?ype=20terms?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 38 +++++++++++++++++-- 1 file changed, 35 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index d511625d9..c3610ada5 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -6,8 +6,23 @@ This document is the single source of truth for the schema, and column design fo Unlike existing Database engines, which work on raw data or explicitly defined materialized tables with schema and column names provided by the users, ASAPPlanner is designed for querying and execution over the mix of raw data and ASAP Primitives. ASAP primitives are usually compact summaries over raw data. Therefore, it introduces new requirement when we design the schema and node definitions for LogicalASAPDAG and PhysicalASAPDAG. -Assuming we have the Logical DAG defined for a canonicalized representation for a batch of queries. [TODO: add links for this here. ] -The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG [TODO: link PR 511's doc here], but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives: SummaryCreation?, SummaryUpdate, SummaryMerge, SummaryDelete, SummarySubtraction, SummaryEstimate [TODO: check what is the complete list or discuss with others about the list]. +Assuming we have the Logical DAG defined for a canonicalized representation for a batch of queries (the `LogicalDAG` produced by the frontends, [planning stages §0](planner-layering.md#0-language-specific-frontends); the stages that follow are in [planner-layering.md](planner-layering.md#stages)). +The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG ([decoupling operators from scalar expressions](decoupling_op_and_expr.md), with the unified operator type in [operator sharing §1.1](operator-sharing.md#11-unified-operator-type)), but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: + +| Operator | Input → output | Status | +|---|---|---| +| `SummaryAgg` | values → summary state (creates and updates the state; `SummaryUpdate` is its input mapping, not a separate operator) | implemented | +| `SummaryEstimate` | sketch state → value | implemented | +| `FinalizeExactAccumulator` | exact accumulator state → value | implemented | +| `MaintainPopulation` | values → maintained membership (state) | implemented | +| `EvaluatePopulation` | maintained membership → value | implemented | +| `SummaryMerge` | state × N → state | structure in #560, coverage check in #646 | +| `SummarySubtract` | state × state → state | reserved | +| `SummaryDelete` | state → state without one key | reserved | +| `SummaryJoin` | state × state → state | reserved | +| `Extension` | state → state, named by an extension | reserved | + +§5 walks through each of them. Each of the Summary operators also require the ASAP primitive information above to inter-operate correctly, preserving semantic correctness. Basically, the following information should be represented to preserve the equivalent query semantics when we introduce ASAP Primitives to logical query representation, and following physical one. @@ -27,13 +42,30 @@ Therefore, these requirements drive the following schema and metadata, node info ## 2. Existing database terminology for schema, table, column, and physical data layout +This section fixes the words used below. They follow relational databases and Apache Arrow / DataFusion, which ASAPPlanner's frontend already uses. + +| Term | Meaning in existing systems | In ASAPPlanner | +|---|---|---| +| **Relation / table** | A set (bag) of rows with the same columns. A base table is stored; a derived relation is the output of a query operator. | Every edge in the DAG carries a relation. A `Scan` reads a base table (SQL table or PromQL metric); every other operator outputs a derived relation. | +| **Row / tuple** | One element of a relation: one value per column. | One output row of a node. For PromQL, one sample of one series at one time. | +| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.3). | +| **Schema** | The ordered list of columns of a relation: name, data type, nullability (Arrow `Schema` of `Field { name, data_type, nullable }`; DataFusion `DFSchema` adds the table qualifier). The schema is *metadata*: it describes rows, it contains none. | `Schema` of `Field { name, dtype, nullable, table }` in `crates/types/src/pre_asap/schema.rs`. Unlike Arrow, `dtype` can be a summary state type (§3). | +| **Data type** | The type of a column's values (`Int64`, `Utf8`, `Timestamp`, …). | `DataType`, wrapped as `FieldDataType::Plain`. | +| **Aggregate state** | The intermediate value of an aggregate function before its final result, e.g. `(sum, count)` for `AVG` (DataFusion `Accumulator::state`, partial/final aggregation). It is never exposed as a column type to users. | Summary state *is* a column type here (`FieldDataType::Sketch`, `ExactAggregate`, …), so state can flow along edges and be merged, stored and read by later operators. | +| **View / materialized view** | A view is a named query (its *definition*). A materialized view also stores the query's result rows; a query can then be answered from it when its definition matches (view matching, §4.1). | A built summary state is a materialized aggregation view whose aggregate is a summary family. Its definition and which rows it took are its coverage (§4). | +| **Physical data layout** | How rows are stored: row-oriented or columnar (Arrow `RecordBatch`: one array per column), split into partitions (hash or range) and batches. | Decided in physical planning ([planning stages §2](planner-layering.md#2-physical-asap-aware-optimization)) and by the executing backend. The logical schema does not depend on it. | + +Two consequences for the design: + +- A schema says what *kind* of values flow along an edge, never *which* rows. Which rows a relation contains is decided by the operators below it (its definition). This is why coverage is a node property and not part of the schema (§4). +- Existing systems keep aggregate state internal to one operator. ASAPPlanner makes it a first-class column type so that one state can be shared, merged and stored across queries, which is what §3 and §4 add. ## 3. Proposed schema design Schema represents the **metadata** of information flow along an **edge** between two nodes in a logical or physical DAG. The schema field is associated with the node in the DAG. The consumer of the node in the DAG takes the schema from the producer node as input. Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. Based on our requirement, each field should contain the following information. -1. **What type of the ASAP Primitive is** A state column can be a raw data type (e.g., numerical number, string). It can also be a [summary type](TODO: add link), e.g., the summary family is sketch, and the sketch type is quantile KLL sketch algorithm, and KLL sketch has K as parameter as the schema. (TODO: confirm the terminology with corresponding code/doc) It has a family, an algorithm and parameters. +1. **What type of the ASAP Primitive is** A state column can be a raw data type (e.g., numerical number, string). It can also be a [summary type](#61-schema-and-field-types-cratestypessrcpre_asapschemars) (`FieldDataType` in `crates/types/src/pre_asap/schema.rs`), e.g., the summary family is sketch, and the sketch type is quantile KLL sketch algorithm, and KLL sketch has K as parameter as the schema. In the code these are: **family** = the `FieldDataType` variant (`ExactAggregate`, `Sketch`, `Sample`, `Wavelet`, `StatModel`; `Plain` is a raw value); for sketches, **category** = `SketchCategory` (`Quantile`, `Frequency`, `Cardinality`, `TopK`, `Universal`), **algorithm** = `SketchAlgorithm` (`Kll`, `Cms`, `Hll`, …) and **parameters** = `SketchParams` (`Kll { k }`), bundled as `SketchKind` ([§6.2](#62-state-family-parameters-cratestypessrcpost_asapsketchrs)); a sketch also carries its `GroupingStrategy`. So the example is `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, …)`. It has a family, an algorithm and parameters. 2. **What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents** This information is being mapped based on the primitive type. From 47882b7ce466f0cadec2b3f58c1898009eaec333 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 8 Oct 2026 15:44:47 +0000 Subject: [PATCH 11/59] docs: point links into asap-primitive-schema.md at its current sections Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/physical-planning-and-deployment.md | 2 +- docs/design_docs/proposals/decoupling_op_and_expr.md | 2 +- docs/develop_docs/asap-aware-mapping-contracts.md | 2 +- docs/develop_docs/pre-asap-ir.md | 2 +- 4 files changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/design_docs/physical-planning-and-deployment.md b/docs/design_docs/physical-planning-and-deployment.md index c48dfed54..00caf5d51 100644 --- a/docs/design_docs/physical-planning-and-deployment.md +++ b/docs/design_docs/physical-planning-and-deployment.md @@ -122,7 +122,7 @@ a complete summary computation: the same four fields can summarize different value expressions or produce different states. The semantic information a summary depends on, and where the IR records each part (field type, producing operator, or coverage), is specified in -[Schema and physical data for ASAP primitives](proposals/asap-primitive-schema.md#23-consideration-3-the-metadata-preserves-summary-semantics). +[Schema and physical data for ASAP primitives](proposals/asap-primitive-schema.md#4-proposed-node-field-design). The canonical selected computation is authoritative. Those categories describe what must be preserved, not a new flat IR or a second expression language. diff --git a/docs/design_docs/proposals/decoupling_op_and_expr.md b/docs/design_docs/proposals/decoupling_op_and_expr.md index f850116bb..962af6c5c 100644 --- a/docs/design_docs/proposals/decoupling_op_and_expr.md +++ b/docs/design_docs/proposals/decoupling_op_and_expr.md @@ -62,7 +62,7 @@ operator inputs and scalar query-result references use `Rc`. read by expressions. Keep `ScalarExpr::Column(ColumnId)`: the ID selects a field for type checking and the corresponding input value for evaluation, independently of the executor's row/column storage layout. See the -[fields versus column references contract](asap-primitive-schema.md#21-consideration-1-the-schema-is-the-edge-between-two-nodes). +[fields versus column references contract](asap-primitive-schema.md#3-proposed-schema-design). Names are resolved to `ColumnId` before constructing these nodes. Parsing and unresolved `ColumnRef` handling remain frontend concerns; no alternative generic diff --git a/docs/develop_docs/asap-aware-mapping-contracts.md b/docs/develop_docs/asap-aware-mapping-contracts.md index 011758143..0e16d4116 100644 --- a/docs/develop_docs/asap-aware-mapping-contracts.md +++ b/docs/develop_docs/asap-aware-mapping-contracts.md @@ -328,7 +328,7 @@ backend inspection. Automatic selection skips those unproven ratios. Use A summary's identity has four levels: family (`FieldDataType` variant), sketch category (`SketchCategory`), algorithm (`SketchAlgorithm`), and the validated committed choice (`SketchKind`). The levels and their validation are specified in -[Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#22-consideration-2-a-field-can-have-an-asap-primitive-type). +[Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#3-proposed-schema-design). Where this matters in practice: `CostModel::rank_candidates`, `CostModel::size_params`, and `SketchAlgorithmStrategy::replacements` operate at the **algorithm** level. `summary_candidates(intent)` returns a list of `SketchAlgorithm`s (`[Kll, DDSketch]` for a `Quantile` intent), never a bare `SketchKind` with nothing chosen underneath it. `SketchKind` appears after an algorithm has been selected and sized—on `Realization::Sketch(SketchKind)` and `FieldDataType::Sketch(SketchKind, GroupingStrategy)`. diff --git a/docs/develop_docs/pre-asap-ir.md b/docs/develop_docs/pre-asap-ir.md index 2492e28d7..77d80c847 100644 --- a/docs/develop_docs/pre-asap-ir.md +++ b/docs/develop_docs/pre-asap-ir.md @@ -20,7 +20,7 @@ The pre-ASAP IR is defined using the `QueryExpr` enum. We discuss some of import `Schema` holds `Field` metadata (name, type, nullability, qualifier) and no values; an unresolved `ColumnRef` resolves to a positional `ColumnId` within one schema. The design, including how the same position selects a runtime value, is -in [Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#21-consideration-1-the-schema-is-the-edge-between-two-nodes). +in [Schema and physical data for ASAP primitives](../design_docs/proposals/asap-primitive-schema.md#3-proposed-schema-design). ## Node index From 0ba9b2131b825170a61f8e3add7f95aa9c96f4e1 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 17:49:24 +0000 Subject: [PATCH 12/59] docs: list the levels of a summary type as a table Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 16 +++++++++++++++- 1 file changed, 15 insertions(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index c3610ada5..ea9df1203 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -65,7 +65,21 @@ Schema represents the **metadata** of information flow along an **edge** between Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. Based on our requirement, each field should contain the following information. -1. **What type of the ASAP Primitive is** A state column can be a raw data type (e.g., numerical number, string). It can also be a [summary type](#61-schema-and-field-types-cratestypessrcpre_asapschemars) (`FieldDataType` in `crates/types/src/pre_asap/schema.rs`), e.g., the summary family is sketch, and the sketch type is quantile KLL sketch algorithm, and KLL sketch has K as parameter as the schema. In the code these are: **family** = the `FieldDataType` variant (`ExactAggregate`, `Sketch`, `Sample`, `Wavelet`, `StatModel`; `Plain` is a raw value); for sketches, **category** = `SketchCategory` (`Quantile`, `Frequency`, `Cardinality`, `TopK`, `Universal`), **algorithm** = `SketchAlgorithm` (`Kll`, `Cms`, `Hll`, …) and **parameters** = `SketchParams` (`Kll { k }`), bundled as `SketchKind` ([§6.2](#62-state-family-parameters-cratestypessrcpost_asapsketchrs)); a sketch also carries its `GroupingStrategy`. So the example is `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, …)`. It has a family, an algorithm and parameters. +1. **What type of the ASAP Primitive is.** The field's type is a [`FieldDataType`](#61-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: + + - **Raw value**: `Plain(DataType)`, e.g., a number or a string. + - **Summary state**: described from coarse to fine by four levels: + + | Level | Code | Values | KLL example | + |---|---|---|---| + | Family | `FieldDataType` variant | `ExactAggregate`, `Sketch`, `Sample`, `Wavelet`, `StatModel` | `Sketch` | + | Category (sketches only) | `SketchCategory` | `Quantile`, `Frequency`, `Cardinality`, `TopK`, `Universal` | `Quantile` | + | Algorithm | `SketchAlgorithm` (other families: `ExactKind`, `SamplingKind`, …) | `Kll`, `Cms`, `Hll`, `DDSketch`, … | `Kll` | + | Parameters | `SketchParams` (other families: `ExactParams`, `SamplingParams`, …) | per algorithm | `Kll { k: 200 }` | + + - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.2](#62-state-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. + - So a quantile KLL sketch with `k = 200`, one instance per group, has the type `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, PerSubpopulationInstance)`. + 2. **What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents** This information is being mapped based on the primitive type. From 03cea00a5f84f23b683a94d1e492519794e419a8 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 17:50:25 +0000 Subject: [PATCH 13/59] docs: list supported query intents per summary type Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 21 ++++++++++++++++++- 1 file changed, 20 insertions(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index ea9df1203..032cd378c 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -80,7 +80,26 @@ Based on our requirement, each field should contain the following information. - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.2](#62-state-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. - So a quantile KLL sketch with `k = 200`, one instance per group, has the type `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, PerSubpopulationInstance)`. -2. **What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents** This information is being mapped based on the primitive type. +2. **What query intent the summarized ASAP Primitive can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: + + - **Statistical aggregation intents** (`AggIntent`, `crates/types/src/pre_asap/agg_intent.rs`): which aggregate the state can answer, and how it is read out. + + | Intent | Candidate types | Readout | + |---|---|---| + | `Quantile` | `Sketch`: `Kll`, `DDSketch` | `SummaryEstimate(Quantile { q })` | + | `Cardinality` | `Sketch`: `Hll`, `Theta`, `Kmv`, `UnivMon` (one column only) | `SummaryEstimate(Cardinality)` | + | `Count` (approximate) | `Sketch`: `Cms`, `CountSketch`, `UnivMon` | `SummaryEstimate(PointCount { .. })` | + | `TopK` | `Sketch`: `CmsWithHeap`, `CountSketchWithHeap` | `SummaryEstimate(TopK { k })` | + | `FrequencyL2`, `FrequencyEntropy` | `Sketch`: `UnivMon` | `SummaryEstimate(FrequencyL2 \| FrequencyEntropy)` | + | `Sum`, `Count`, `Min`, `Max`, `Rate`, `Increase` (exact) | `ExactAggregate(ExactKind, …)` | `FinalizeExactAccumulator` | + + The sketch candidates are `summary_candidates(intent)` in `crates/asap-aware-mapping/src/replacement.rs`; the readouts are `SketchStatistic` ([§6.3](#63-update-input-and-readouts-post_asapsketchrs-post_asapmaintained_populationrs)). + + - **Time window aggregation intents**: whether states built over smaller windows can answer a larger one. This depends on how the family combines states: + - **Merge** (`SummaryMerge`, §5.6): states over disjoint panes combine into the state of their union, e.g. two 1-minute KLL states answer a 2-minute quantile. Requires a mergeable family. + - **Subtract** (`SummarySubtract`, reserved): a sliding window as a larger state minus an older one. Only families with an inverse (e.g. exact `Sum`/`Count`, CMS) can do this. + + Which inputs may be merged is decided by coverage (§4), not by the schema. ## 4. Proposed Node field design From 883c6efff097aa3d8db2ef00327968d62b0ca18e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 17:57:28 +0000 Subject: [PATCH 14/59] =?UTF-8?q?docs:=20structure=20=C2=A74=20with=20summ?= =?UTF-8?q?ary=20table,=20rule=20tables=20and=20bullet=20lists?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 145 +++++++++++++----- 1 file changed, 106 insertions(+), 39 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 032cd378c..72ec27326 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -104,18 +104,35 @@ Based on our requirement, each field should contain the following information. ## 4. Proposed Node field design -A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. +A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. This field is the node's **coverage**. -**Why coverage is not part of the schema.** Two summary states worth merging always cover different data. `SummaryMerge` requires all inputs to have the same schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2: +**At a glance** -| | State A | State B | Equal? | -|---|---|---|---| -| schema | `(job: Utf8, state: KLL{k=200})` | `(job: Utf8, state: KLL{k=200})` | yes, so the merge is allowed | -| what it summarizes | time `[0,1)` | time `[1,2)` | no, which is why merging them is useful | +| Question | Answer | Section | +|---|---|---| +| Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | +| What is it based on? | Goldstein & Larson view matching (SIGMOD 2001) | §4.1 | +| What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | +| Who sets it? | Nobody: it is derived from the sub-DAG | §4.3 | +| What uses it? | merge, rollup, slice, reuse, subtract | §4.4 | +| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.5 | +| What is left out? | source dependencies, readiness, absolute time binding | §4.6 | + +**Why coverage is not part of the schema.** + +- `SummaryMerge` requires all inputs to have the same schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). +- Two summary states worth merging always cover different data. For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2: + + | | State A | State B | Equal? | + |---|---|---|---| + | schema | `(job: Utf8, state: KLL{k=200})` | `(job: Utf8, state: KLL{k=200})` | yes, so the merge is allowed | + | what it summarizes | time `[0,1)` | time `[1,2)` | no, which is why merging them is useful | + +- If coverage were part of the schema, these two schemas would differ and the merge would be rejected. The only merge left would be a state with an exact copy of itself, which counts every observation twice. -If what a state summarizes were part of the schema, these two schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. So the schema says *what kind of state* this is, and coverage says *which data it was built from*. +So the schema says *what kind of state* this is, and coverage says *which data it was built from*. -A summary state summarizes the result of a whole computation, not a few columns of a raw table. A KLL over `rate(requests_total[5m])` summarizes rate outputs, and a KLL over a join summarizes join rows. So coverage describes the state by the sub-DAG below it, split into the part that says *what is computed* and the part that says *which of its rows were taken*. +**Why coverage is a sub-DAG, not a few table columns.** A summary state summarizes the result of a whole computation: a KLL over `rate(requests_total[5m])` summarizes rate outputs, and a KLL over a join summarizes join rows. So coverage describes the state by the sub-DAG below it, split into *what is computed* and *which of its rows were taken*. ### 4.1 Design basis: view matching (Goldstein & Larson) @@ -123,9 +140,22 @@ The design follows the view matching algorithm of Goldstein and Larson, which de > J. Goldstein and P.-Å. Larson. *Optimizing Queries Using Materialized Views: A Practical, Scalable Solution.* SIGMOD 2001. -The algorithm splits a view's `WHERE` into column equivalence classes, a **range** per column and **residual** predicates. A view can answer a query when the residuals match, the query's ranges lie inside the view's (§3.1.2), the columns needed by compensating predicates are in the view output (§3.3, requirement 2), and the query's `GROUP BY` is a subset of the view's, so the query's groups are further aggregations of the view's groups (§3.3, requirement 3). The SPJ part is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`); it rejects `Aggregate` and `Join` input plans. +**The algorithm.** It splits a view's `WHERE` into three parts: -A summary state is an aggregation view whose aggregate is a summary family. The mapping is: +- column **equivalence classes**, +- a **range** per column, +- **residual** predicates (everything else). + +A view can answer a query when all of these hold: + +1. the residuals match; +2. the query's ranges lie inside the view's (§3.1.2); +3. the columns needed by compensating predicates are in the view output (§3.3, requirement 2); +4. the query's `GROUP BY` is a subset of the view's, so the query's groups are further aggregations of the view's groups (§3.3, requirement 3). + +**Existing implementation.** The SPJ part is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`). It rejects `Aggregate` and `Join` input plans. + +**Mapping.** A summary state is an aggregation view whose aggregate is a summary family: | Goldstein & Larson | Summary coverage | |---|---| @@ -136,7 +166,7 @@ A summary state is an aggregation view whose aggregate is a summary family. The | compensating predicate on view output | slice on a column of `G` only (§4.4) | | query `GROUP BY` ⊆ view `GROUP BY` | rollup (§4.4) | -What this design adds beyond the paper: +**What this design adds beyond the paper:** - **Unions of states.** The paper considers single-view substitutes and notes that requirement 1 "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. - **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.4). @@ -155,12 +185,14 @@ state_g = family( input( σ( C ) ) ) for each group value g of G, restricted to Coverage stores exactly these two things: -- **`definition`**: the `SummaryAgg` node itself, with the selection removed from its child sub-DAG. It carries `C`, `input`, `family` and `G`. It is what the state *means*. -- **`selection`**: a union of boxes over the output columns of `C`. It is *which rows* the state took. +| Part | Contents | Meaning | +|---|---|---| +| **`definition`** | the `SummaryAgg` node itself, with the selection removed from its child sub-DAG; carries `C`, `input`, `family` and `G` | what the state *means* | +| **`selection`** | a union of boxes over the output columns of `C` | *which rows* the state took | -If two states have the same `definition`, their contributions come from the same rows of the same computation, whatever `C` contains (join, union, `rate`, dedup). Disjoint selections then cannot share a row, so no observation is counted twice. No per-operator occurrence rule is needed. +**Why this is enough.** If two states have the same `definition`, their contributions come from the same rows of the same computation, whatever `C` contains (join, union, `rate`, dedup). Disjoint selections then cannot share a row, so no observation is counted twice. No per-operator occurrence rule is needed. -Examples of what ends up where: +**Examples** of what ends up where: | Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | |---|---|---| @@ -174,46 +206,75 @@ In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `r ### 4.3 Deriving the selection -Coverage is derived from the node, never declared. Walking down from the `SummaryAgg` (its own `filter` included), a predicate conjunct goes into `selection` when both hold: +Coverage is derived from the node, never declared. Walking down from the `SummaryAgg` (its own `filter` included), a predicate conjunct goes into `selection` when both rules hold. Otherwise it stays in `definition` as a residual, as in Goldstein & Larson. -1. **It can be lifted to the `SummaryAgg`.** Lifting is the inverse of DataFusion's `PushDownFilter` (`datafusion-optimizer`, `push_down_filter.rs`): a predicate passes `Filter`, `TimeRange`/`TimeShift` and a direct-column `Project` (renaming the column); passes an `Aggregate` only when every column it uses is a group column; and passes a window function or a per-series temporal function such as `rate` only when every column it uses is a partition column (a series label). Anything else stops it. -2. **It is one of the box constraints.** Per column, one of: - - **value set**: `In` or `NotIn` a set of literals, from `=`, `!=`, `IN`, `NOT IN` and `OR` of equalities (as DataFusion's `LiteralGuarantee` extracts them); - - **interval**: lower and upper `std::ops::Bound` (`Included`, `Excluded` or `Unbounded`) from comparisons (as DataFusion's `Interval`); - - **hash partition**: `hash(columns) mod n = k`. +**Rule 1: it can be lifted to the `SummaryAgg`.** Lifting is the inverse of DataFusion's `PushDownFilter` (`datafusion-optimizer`, `push_down_filter.rs`): -A conjunct that fails either rule stays in `definition` as a residual, as in Goldstein & Larson. Column equalities (`a = b`) are residuals too: there are no column equivalence classes. +| Operator on the path | The predicate passes when | +|---|---| +| `Filter`, `Scan.predicates`, `SummaryAgg.filter` | always | +| range `TimeRange`, `TimeShift` without `@` | always | +| `Project` | the column is a direct column item (renaming keeps its identity) | +| `Aggregate` (later) | every column it uses is a group column | +| window function, per-series temporal function such as `rate` (later) | every column it uses is a partition column (a series label) | +| anything else | never | + +**Rule 2: it is a box constraint on one column.** + +| Constraint | From | DataFusion analogue | +|---|---|---| +| value set: `In` / `NotIn` literals | `=`, `!=`, `IN`, `NOT IN`, `OR` of equalities | `LiteralGuarantee` | +| interval: lower and upper `std::ops::Bound` (`Included`, `Excluded`, `Unbounded`) | `<`, `<=`, `>`, `>=` | `Interval` | +| hash partition (later): `hash(columns) mod n = k` | partitioned producers | `Partitioning::Hash` | + +Column equalities (`a = b`) are residuals: there are no column equivalence classes. -Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns, and a direct alias keeps the identity of the column it renames. A column whose `(table, name)` is not unique in the output (two items aliased `k`) cannot be named, so its conjuncts stay residual. Value sets compare literals by type: `1` and `1.0` are never proven different. +**Column identity.** + +- Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns. +- A direct alias keeps the identity of the column it renames. +- A column whose `(table, name)` is not unique in the output (two items aliased `k`) cannot be named, so its conjuncts stay residual. +- Value sets compare literals by type: `1` and `1.0` are never proven different. **Time** is a selection like any other: -- **Absolute** time needs nothing special: it is an interval on the timestamp column (the schema's `time_index`), for example `ts >= t0 AND ts < t1` gives `(Included(t0), Excluded(t1))` on `ts`. The IR has no timestamp literal yet, so such SQL filters stay residual until it does. -- **Relative** time has its own field, because it is not a column value: a `TimeRange(w)` over a `TimeShift(s)` on the lifted chain gives `(Excluded(−(s+w)), Included(−s))` relative to evaluation. PromQL ranges are left-open, matching the executor (`series_window.rs`). This is how Stage 2 tumbling panes are built (`window_composition.rs` in #601), so their time is derived rather than declared. Time is lifted only from a single range `TimeRange`; an instant `TimeRange` picks the latest sample per series, which is not a selection of rows, so it stays in `definition`. +| Kind | How it is selected | Example | +|---|---|---| +| **Absolute** | an interval on the timestamp column (the schema's `time_index`) | `ts >= t0 AND ts < t1` → `(Included(t0), Excluded(t1))` on `ts` | +| **Relative** | its own field `relative_time`, because it is not a column value: one range `TimeRange(w)` over a `TimeShift(s)` | `(Excluded(−(s+w)), Included(−s))` relative to evaluation | -Relative time and a timestamp-column interval are different dimensions, so they are never compared: two states restricted only by different kinds of time are treated as possibly overlapping. Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. +- PromQL ranges are left-open, matching the executor (`series_window.rs`). +- The IR has no timestamp literal yet, so absolute SQL time filters stay residual until it does. +- Stage 2 tumbling panes (`window_composition.rs` in #601) get their time this way, so it is derived rather than declared. +- An instant `TimeRange` picks the latest sample per series, which is not a selection of rows, so it stays in `definition`. +- Relative time and a timestamp-column interval are different dimensions, so they are never compared: two states restricted only by different kinds of time are treated as possibly overlapping. +- Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. ### 4.4 Operations | Operation | Example | Valid when | |---|---|---| | merge (`SummaryMerge`, same `G`) | `[0,1m)` ⊕ `[1m,2m)`; `region='us'` ⊕ `region='eu'` | all `definition`s equal; selections related as the family requires (below) | -| rollup (`SummaryMerge` with `group_by: G'`) | `by[region, job]` → `by[job]` | `G'` ⊆ `G` and the family merges. Groups of one state are disjoint because a row has one value per group column, so no selection check is needed | +| rollup (`SummaryMerge` with `group_by: G'`, later) | `by[region, job]` → `by[job]` | `G'` ⊆ `G` and the family merges. Groups of one state are disjoint because a row has one value per group column, so no selection check is needed | | slice | `by[region, job]` state answering `region = 'us' … by[job]` | the restricted columns are all in `G`. A sketch cannot be filtered, so a restriction on any other column is invalid | | reuse for a query | a stored state answers a query | Goldstein & Larson containment: same `definition`, query selection inside the state's, any compensating restriction is a slice | | subtract (`SummarySubtract`, reserved) | `[0,10) − [0,5)` | same `definition`; the right selection is contained in the left | -One `SummaryMerge { children, group_by }` covers both merge and rollup: one child with a coarser `group_by` is a rollup, and `group_by` equal to the children's is a plain merge. Its coverage is the children's `definition` with `G'` and the union of their selections; adjacent intervals are joined, gaps stay as separate boxes. +**Merge and rollup are one operator.** `SummaryMerge { children, group_by }`: one child with a coarser `group_by` is a rollup, and `group_by` equal to the children's is a plain merge. Its coverage is the children's `definition` with `G'` and the union of their selections; adjacent intervals are joined, gaps stay as separate boxes. -How selections must relate is declared by the family, next to whether it merges (`FieldDataType::family_merges`, added in #592): +**How selections must relate** is declared by the family, next to whether it merges (`FieldDataType::family_merges` and `merge_relation`, added in #592): -- **disjoint** for counting families (KLL, Count-Min, exact `Sum`/`Count`): an overlapping row would be counted twice; -- **overlap allowed** for idempotent families (HLL, exact `Min`/`Max`, distinct sets); -- **contained** for subtraction. +| Relation | Families | Why | +|---|---|---| +| **disjoint** | counting families: KLL, Count-Min, exact `Sum`/`Count` | an overlapping row would be counted twice | +| **overlap allowed** | idempotent families: HLL, exact `Min`/`Max`, distinct sets | adding a row twice does not change the state | +| **contained** | subtraction | the right state must be part of the left | -Two `definition`s are equal when their canonical forms (`canonicalize`) are structurally equal, ignoring planning metadata: `timing`, `guarantee` and `coverage_cache`. `SummaryUpdate.weight_domain` is compared: it is derived from `C` and `input`, so it differs only if a derivation is wrong. A state built at ingestion time and one built at query time can therefore merge. +**When two `definition`s are equal.** -States over different sources have different `definition`s and do not merge. To combine tables, put `UNION ALL` with a marker column below one `SummaryAgg`; the marker is then an ordinary column for `selection` or `G`. +- They must be structurally equal, ignoring planning metadata: `timing`, `guarantee` and `coverage_cache`. So a state built at ingestion time and one built at query time can merge. +- `SummaryUpdate.weight_domain` is compared: it is derived from `C` and `input`, so it differs only if a derivation is wrong. +- States over different sources have different `definition`s and do not merge. To combine tables, put `UNION ALL` with a marker column below one `SummaryAgg`; the marker is then an ordinary column for `selection` or `G`. ### 4.5 Interface @@ -264,15 +325,21 @@ impl SummaryCoverage { } ``` -`OperatorNode::new` still rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. +- `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. +- A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. ### 4.6 What coverage does not contain -- **Source dependencies**: which source rows must be read to compute the contributions, such as the 5-minute window under `rate`. This is read planning and maintenance (compare `materialized/dependencies.rs` in `datafusion-materialized-views`). -- **Readiness and completeness**: whether a stored instance holds all of its rows. -- **Absolute binding** of relative time, and deployment identity. +| Left out | Example | Owner | +|---|---|---| +| **Source dependencies**: which source rows must be read to compute the contributions | the 5-minute window under `rate` | ASAPQuery-backend: read planning and maintenance (compare `materialized/dependencies.rs` in `datafusion-materialized-views`) | +| **Readiness and completeness**: whether a stored instance holds all of its rows | a pane still being filled | ASAPQuery-backend | +| **Absolute binding** of relative time, and deployment identity | pane `(−3m, −2m]` at evaluation time `t` | ASAPQuery-backend | + +**SDS mapping.** The SDS split matches coverage: -These belong to ASAPQuery-backend. The SDS split matches coverage: `SummaryDefinition` stores the serialized `definition` (Planner provides its serde; Backend owns the format version, definition id and hash), and a `StoredSummary`'s coordinates are the `selection` bound to one evaluation plus the group value. +- `SummaryDefinition` stores the serialized `definition`. Planner provides its serde; Backend owns the format version, definition id and hash. +- A `StoredSummary`'s coordinates are the `selection` bound to one evaluation, plus the group value. ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators From 7cac25418c3bbb8c994f73a4909854780e32aa62 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 17:57:59 +0000 Subject: [PATCH 15/59] =?UTF-8?q?docs:=20rename=20=C2=A74.2=20heading?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 72ec27326..f36c49af4 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -172,7 +172,7 @@ A view can answer a query when all of these hold: - **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.4). - **Value sets and hash partitions** next to ranges, and **evaluation-relative time** (§4.3). -### 4.2 Coverage = definition + selection +### 4.2 Summary Coverage = Summary definition + selection A state built by `SummaryAgg` means From 36dc7f2e5ace16b16fad3cdad7d7e11de067a520 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:00:00 +0000 Subject: [PATCH 16/59] docs: explain the residual condition of view matching Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index f36c49af4..9122293e9 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -148,7 +148,7 @@ The design follows the view matching algorithm of Goldstein and Larson, which de A view can answer a query when all of these hold: -1. the residuals match; +1. every residual predicate of the view also appears in the query (§3.1.2, residual subsumption, checked by matching the predicates' text after normalization). Residuals cannot be reasoned about, so the view must not filter out any row the query needs. For example, a view with `WHERE lower(name) LIKE 'a%'` can answer a query only if the query has the same `lower(name) LIKE 'a%'`. The query may have extra residuals; they are applied to the view's output as compensating predicates; 2. the query's ranges lie inside the view's (§3.1.2); 3. the columns needed by compensating predicates are in the view output (§3.3, requirement 2); 4. the query's `GROUP BY` is a subset of the view's, so the query's groups are further aggregations of the view's groups (§3.3, requirement 3). From 25c91745269cdf164aee4b0d18234464531cd563 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:00:08 +0000 Subject: [PATCH 17/59] docs: define the three parts of a view's WHERE with examples Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 9122293e9..47d91f6cc 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -142,9 +142,9 @@ The design follows the view matching algorithm of Goldstein and Larson, which de **The algorithm.** It splits a view's `WHERE` into three parts: -- column **equivalence classes**, -- a **range** per column, -- **residual** predicates (everything else). +- column **equivalence classes**: columns known to be equal, from column equalities such as `a = b` (e.g. join keys); +- a **range** per column: the interval a column must lie in, from comparisons with a constant such as `x > 5 AND x <= 10` (an equality `x = 5` is the range `[5, 5]`); +- **residual** predicates: every conjunct that is neither of the above, such as `a + b > 10`, `lower(name) LIKE 'a%'`, or `x = 1 OR y = 2`. The algorithm does not interpret them; it only checks whether the same predicate appears in the query. A view can answer a query when all of these hold: From a90b667ccbaf93d0efe086f6109b36af0aae29ee Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:00:30 +0000 Subject: [PATCH 18/59] docs: define equivalence classes, ranges and residuals Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 47d91f6cc..f012c0683 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -142,9 +142,9 @@ The design follows the view matching algorithm of Goldstein and Larson, which de **The algorithm.** It splits a view's `WHERE` into three parts: -- column **equivalence classes**: columns known to be equal, from column equalities such as `a = b` (e.g. join keys); -- a **range** per column: the interval a column must lie in, from comparisons with a constant such as `x > 5 AND x <= 10` (an equality `x = 5` is the range `[5, 5]`); -- **residual** predicates: every conjunct that is neither of the above, such as `a + b > 10`, `lower(name) LIKE 'a%'`, or `x = 1 OR y = 2`. The algorithm does not interpret them; it only checks whether the same predicate appears in the query. +- **Column equivalence classes.** An equivalence class is a set of columns that have the same value in every row that satisfies the `WHERE`. They come from column equalities: `a = b AND b = c` gives the class `{a, b, c}`. For example, the join condition `orders.cust_id = customers.id` puts both columns in one class. The algorithm uses the classes to recognize that a query and a view mean the same thing even when they name different but equal columns. +- **A range per column.** The range of a column (more precisely, of an equivalence class) is the interval of values it can have in rows that satisfy the `WHERE`. It comes from comparisons with a constant: `x > 5 AND x <= 10` gives `x ∈ (5, 10]`, and `x = 5` gives `x ∈ [5, 5]`. A column with no such comparison has the range `(−∞, +∞)`. Ranges let the algorithm prove containment: a view with `x ∈ (0, 100]` holds every row a query with `x ∈ (5, 10]` needs. +- **Residual predicates.** A residual predicate is a conjunct of the `WHERE` (one of its `AND`-ed terms) that is neither a column equality nor a comparison of a column with a constant, such as `a + b > 10`, `lower(name) LIKE 'a%'`, or `x = 1 OR y = 2`. The algorithm does not interpret them; it only checks whether the same predicate appears in the query. A view can answer a query when all of these hold: From 3cecd672540bc6f5f3c9fb7d0866b989f9c90338 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:10:08 +0000 Subject: [PATCH 19/59] =?UTF-8?q?docs:=20move=20the=20Goldstein=20&=20Lars?= =?UTF-8?q?on=20mapping=20after=20=C2=A74.2=20as=20=C2=A74.3?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 65 ++++++++++--------- 1 file changed, 34 insertions(+), 31 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index f012c0683..0f96569a7 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -48,7 +48,7 @@ This section fixes the words used below. They follow relational databases and Ap |---|---|---| | **Relation / table** | A set (bag) of rows with the same columns. A base table is stored; a derived relation is the output of a query operator. | Every edge in the DAG carries a relation. A `Scan` reads a base table (SQL table or PromQL metric); every other operator outputs a derived relation. | | **Row / tuple** | One element of a relation: one value per column. | One output row of a node. For PromQL, one sample of one series at one time. | -| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.3). | +| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.4). | | **Schema** | The ordered list of columns of a relation: name, data type, nullability (Arrow `Schema` of `Field { name, data_type, nullable }`; DataFusion `DFSchema` adds the table qualifier). The schema is *metadata*: it describes rows, it contains none. | `Schema` of `Field { name, dtype, nullable, table }` in `crates/types/src/pre_asap/schema.rs`. Unlike Arrow, `dtype` can be a summary state type (§3). | | **Data type** | The type of a column's values (`Int64`, `Utf8`, `Timestamp`, …). | `DataType`, wrapped as `FieldDataType::Plain`. | | **Aggregate state** | The intermediate value of an aggregate function before its final result, e.g. `(sum, count)` for `AVG` (DataFusion `Accumulator::state`, partial/final aggregation). It is never exposed as a column type to users. | Summary state *is* a column type here (`FieldDataType::Sketch`, `ExactAggregate`, …), so state can flow along edges and be merged, stored and read by later operators. | @@ -113,10 +113,11 @@ A node in the physical data will represent the data or summary instance, so a no | Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | | What is it based on? | Goldstein & Larson view matching (SIGMOD 2001) | §4.1 | | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | -| Who sets it? | Nobody: it is derived from the sub-DAG | §4.3 | -| What uses it? | merge, rollup, slice, reuse, subtract | §4.4 | -| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.5 | -| What is left out? | source dependencies, readiness, absolute time binding | §4.6 | +| How does it map to the paper? | computation, aggregate and `GROUP BY` → `definition`; ranges → `selection` | §4.3 | +| Who sets it? | Nobody: it is derived from the sub-DAG | §4.4 | +| What uses it? | merge, rollup, slice, reuse, subtract | §4.5 | +| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.6 | +| What is left out? | source dependencies, readiness, absolute time binding | §4.7 | **Why coverage is not part of the schema.** @@ -155,23 +156,6 @@ A view can answer a query when all of these hold: **Existing implementation.** The SPJ part is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`). It rejects `Aggregate` and `Join` input plans. -**Mapping.** A summary state is an aggregation view whose aggregate is a summary family: - -| Goldstein & Larson | Summary coverage | -|---|---| -| SPJ part: tables, joins, residual predicates | the computation `C` below the `SummaryAgg` (§4.2), part of `definition` | -| aggregate function and its argument | `family` and `input` (`SummaryUpdate`) of the `SummaryAgg`, part of `definition` | -| `GROUP BY` | the `SummaryAgg` reduction `G`, part of `definition` | -| ranges per column | `selection` (§4.3) | -| compensating predicate on view output | slice on a column of `G` only (§4.4) | -| query `GROUP BY` ⊆ view `GROUP BY` | rollup (§4.4) | - -**What this design adds beyond the paper:** - -- **Unions of states.** The paper considers single-view substitutes and notes that requirement 1 "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. -- **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.4). -- **Value sets and hash partitions** next to ranges, and **evaluation-relative time** (§4.3). - ### 4.2 Summary Coverage = Summary definition + selection A state built by `SummaryAgg` means @@ -202,9 +186,28 @@ Coverage stores exactly these two things: | `Filter(rate > 0, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))`, `rate > 0` as residual | — | | `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | -In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `rate` and changes the rate values, so it is not a selection of output rows. The 5-minute read window is a source dependency, not coverage (§4.6). +In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `rate` and changes the rate values, so it is not a selection of output rows. The 5-minute read window is a source dependency, not coverage (§4.7). + +### 4.3 Mapping to Goldstein & Larson + +A summary state is an aggregation view whose aggregate is a summary family: + +| Goldstein & Larson | Summary coverage | +|---|---| +| SPJ part: tables, joins, residual predicates | the computation `C` below the `SummaryAgg` (§4.2), part of `definition` | +| aggregate function and its argument | `family` and `input` (`SummaryUpdate`) of the `SummaryAgg`, part of `definition` | +| `GROUP BY` | the `SummaryAgg` reduction `G`, part of `definition` | +| ranges per column | `selection` (§4.4) | +| compensating predicate on view output | slice on a column of `G` only (§4.5) | +| query `GROUP BY` ⊆ view `GROUP BY` | rollup (§4.5) | + +**What this design adds beyond the paper**: + +- **Unions of states.** The paper considers single-view substitutes and notes that requirement 1 "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. +- **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.5). +- **Value sets and hash partitions** next to ranges, and **evaluation-relative time** (§4.4). -### 4.3 Deriving the selection +### 4.4 Deriving the selection Coverage is derived from the node, never declared. Walking down from the `SummaryAgg` (its own `filter` included), a predicate conjunct goes into `selection` when both rules hold. Otherwise it stays in `definition` as a residual, as in Goldstein & Larson. @@ -250,7 +253,7 @@ Column equalities (`a = b`) are residuals: there are no column equivalence class - Relative time and a timestamp-column interval are different dimensions, so they are never compared: two states restricted only by different kinds of time are treated as possibly overlapping. - Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. -### 4.4 Operations +### 4.5 Operations | Operation | Example | Valid when | |---|---|---| @@ -276,7 +279,7 @@ Column equalities (`a = b`) are residuals: there are no column equivalence class - `SummaryUpdate.weight_domain` is compared: it is derived from `C` and `input`, so it differs only if a derivation is wrong. - States over different sources have different `definition`s and do not merge. To combine tables, put `UNION ALL` with a marker column below one `SummaryAgg`; the marker is then an ordinary column for `selection` or `G`. -### 4.5 Interface +### 4.6 Interface ```rust pub struct OperatorNode { @@ -328,7 +331,7 @@ impl SummaryCoverage { - `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. - A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. -### 4.6 What coverage does not contain +### 4.7 What coverage does not contain | Left out | Example | Owner | |---|---|---| @@ -367,7 +370,7 @@ SummaryAgg(family = Sketch(KLL{k=200}, PerSubpopulationInstance), - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. - Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **always derived**, never declared (§4.3). Both conjuncts of the `Filter` lift into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. +- Coverage: **always derived**, never declared (§4.4). Both conjuncts of the `Filter` lift into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. - Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. ### 5.2 `SummaryEstimate`: sketch state → value @@ -439,7 +442,7 @@ EvaluatePopulation(evaluation = PopulationStatistic::Quantile { q: 0.99 }) ### 5.6 `SummaryMerge`: state × N → state (merge and rollup) -On `main`, `SummaryMerge { children }` is **reserved**: `is_unimplemented()` returns true, and `output_schema()` and `validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so `OperatorNode::new` fails. #560 enables it for children with identical schemas. This design adds `group_by`, so one operator does both merge and rollup (§4.4): +On `main`, `SummaryMerge { children }` is **reserved**: `is_unimplemented()` returns true, and `output_schema()` and `validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so `OperatorNode::new` fails. #560 enables it for children with identical schemas. This design adds `group_by`, so one operator does both merge and rollup (§4.5): ```rust SummaryMerge { children: Vec, group_by: Reduction } @@ -464,7 +467,7 @@ Scenario C, rollup: one `KLL(latency) by[region, job]` state merged with `group_ - Output schema: the children's schema with the group key fields reduced to `group_by`. - Checks: - at least one child, every child is `State` with exactly one state field; - - all children have equal `definition`s (§4.4), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; + - all children have equal `definition`s (§4.5), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; - `group_by` ⊆ the children's `G`, and the family merges; - the children's selections relate as the family requires: disjoint for KLL, so pane 0 with pane 0 is rejected; overlap is allowed for HLL. - Coverage: **derived**: the shared `definition` with `group_by`, and the union of the children's selections. Adjacent intervals join; gaps stay as separate boxes. Nested merges work because a child merge has coverage like any other summary node. @@ -476,7 +479,7 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | Operator | Fields | Intended edge shape | |---|---|---| -| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.4) | +| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.5) | | `SummaryDelete` | `summary_input, key: ColumnId` | State → State with the entries for `key` removed | | `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | | `Extension` | `child, name` | deployment-named state operator | From a371b246f6139a14cf67156be329cd7cea244b3e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:10:15 +0000 Subject: [PATCH 20/59] docs: say which paper requirement unions of views relax Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 0f96569a7..fc3ebe2ca 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -203,7 +203,7 @@ A summary state is an aggregation view whose aggregate is a summary family: **What this design adds beyond the paper**: -- **Unions of states.** The paper considers single-view substitutes and notes that requirement 1 "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. +- **Unions of states.** The paper considers single-view substitutes and notes that its requirement 1, that the view contains all rows the query needs, "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. - **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.5). - **Value sets and hash partitions** next to ranges, and **evaluation-relative time** (§4.4). From fbd49cb6310391077687ba710d3989f0a8507a36 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:12:08 +0000 Subject: [PATCH 21/59] docs: SummaryMerge is implemented on main since #560 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 17 +++++++++++------ 1 file changed, 11 insertions(+), 6 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index fc3ebe2ca..7313d26b9 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -16,7 +16,7 @@ The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in | `FinalizeExactAccumulator` | exact accumulator state → value | implemented | | `MaintainPopulation` | values → maintained membership (state) | implemented | | `EvaluatePopulation` | maintained membership → value | implemented | -| `SummaryMerge` | state × N → state | structure in #560, coverage check in #646 | +| `SummaryMerge` | state × N → state | implemented (#560: identical schemas); coverage check in #646 | | `SummarySubtract` | state × state → state | reserved | | `SummaryDelete` | state → state without one key | reserved | | `SummaryJoin` | state × state → state | reserved | @@ -442,7 +442,11 @@ EvaluatePopulation(evaluation = PopulationStatistic::Quantile { q: 0.99 }) ### 5.6 `SummaryMerge`: state × N → state (merge and rollup) -On `main`, `SummaryMerge { children }` is **reserved**: `is_unimplemented()` returns true, and `output_schema()` and `validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so `OperatorNode::new` fails. #560 enables it for children with identical schemas. This design adds `group_by`, so one operator does both merge and rollup (§4.5): +Current state: + +- **On `main` (since #560):** `SummaryMerge { children }` is implemented. `validate_inputs()` accepts it when there is at least one child, every child is `State` with exactly one state field, and all children have identical schemas. The output schema is the children's schema. +- **#646 (open):** adds the coverage check. `OperatorNode::new` and `validate_structure` also require equal `definition`s and disjoint selections, and `coverage()` returns the merged coverage. +- **Planned:** `group_by`, so one operator does both merge and rollup (§4.5): ```rust SummaryMerge { children: Vec, group_by: Reduction } @@ -493,7 +497,7 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | `FinalizeExactAccumulator` | `State` (`ExactAggregate`) | source's value kind | no | none | implemented | | `MaintainPopulation` | `Relation` (table) / `InstantVector` (series) | `State` | yes (by kind; fields plain) | none | implemented | | `EvaluatePopulation` | `State` from `MaintainPopulation` | source's value kind | no | none | implemented | -| `SummaryMerge` | `State` × N | `State` | yes | derived: shared definition with `group_by`, union of selections | reserved; enabled by #560, `group_by` added by this design | +| `SummaryMerge` | `State` × N | `State` | yes | derived: shared definition with `group_by`, union of selections | implemented (#560); coverage check in #646; `group_by` planned | | `SummarySubtract` | `State` × 2 | `State` | yes | derived: left selection minus right (planned) | reserved | | `SummaryDelete` | `State` | `State` | yes | — | reserved | | `SummaryJoin` | `State` × 2 | `State` | yes | — | reserved | @@ -665,8 +669,9 @@ pub enum ASAPOp { FinalizeExactAccumulator { child: Rc }, MaintainPopulation { child: Rc, population: MaintainedPopulation }, EvaluatePopulation { child: Rc, evaluation: PopulationStatistic }, - // Reserved on this branch; #560 implements SummaryMerge. + // Implemented since #560 (identical child schemas). SummaryMerge { children: Vec> }, + // Reserved: migrated but unimplemented. SummarySubtract { left: Rc, right: Rc }, SummaryDelete { summary_input: Rc, key: ColumnId }, SummaryJoin { outer: Rc, inner: Rc, key: ColumnId, family: FieldDataType }, @@ -677,9 +682,9 @@ impl ASAPOp { pub fn children(&self) -> Vec<&Rc>; // SummaryAgg includes its filter's subquery nodes pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Self; pub fn kind_name(&self) -> &'static str; - /// Merge, Subtract, Delete, Join, Extension on this branch; #560 removes Merge. + /// Subtract, Delete, Join, Extension. pub fn is_unimplemented(&self) -> bool; - /// SummaryAgg/SummaryJoin `family`; #560 adds SummaryMerge (its inputs' state type). + /// SummaryAgg/SummaryJoin `family`; SummaryMerge: its inputs' state type. pub fn produced_state(&self) -> Option<&FieldDataType>; pub fn output_schema(&self) -> Result; pub fn output_kind(&self) -> OperatorResultKind; From c1eff12d5a366e4566c0ef2b6276e275770f6bea Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:14:16 +0000 Subject: [PATCH 22/59] =?UTF-8?q?docs:=20=C2=A74.7=20leaves=20the=20deploy?= =?UTF-8?q?ment=20runtime,=20such=20as=20SDS,=20out=20of=20coverage?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 10 +++------- 1 file changed, 3 insertions(+), 7 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 7313d26b9..bec130fab 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -117,7 +117,7 @@ A node in the physical data will represent the data or summary instance, so a no | Who sets it? | Nobody: it is derived from the sub-DAG | §4.4 | | What uses it? | merge, rollup, slice, reuse, subtract | §4.5 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.6 | -| What is left out? | source dependencies, readiness, absolute time binding | §4.7 | +| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.7 | **Why coverage is not part of the schema.** @@ -186,7 +186,7 @@ Coverage stores exactly these two things: | `Filter(rate > 0, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))`, `rate > 0` as residual | — | | `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | -In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `rate` and changes the rate values, so it is not a selection of output rows. The 5-minute read window is a source dependency, not coverage (§4.7). +In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `rate` and changes the rate values, so it is not a selection of output rows. Reading those 5 minutes of source data is the runtime's job, not coverage (§4.7). ### 4.3 Mapping to Goldstein & Larson @@ -333,11 +333,7 @@ impl SummaryCoverage { ### 4.7 What coverage does not contain -| Left out | Example | Owner | -|---|---|---| -| **Source dependencies**: which source rows must be read to compute the contributions | the 5-minute window under `rate` | ASAPQuery-backend: read planning and maintenance (compare `materialized/dependencies.rs` in `datafusion-materialized-views`) | -| **Readiness and completeness**: whether a stored instance holds all of its rows | a pane still being filled | ASAPQuery-backend | -| **Absolute binding** of relative time, and deployment identity | pane `(−3m, −2m]` at evaluation time `t` | ASAPQuery-backend | +Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to ASAPQuery-backend, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. **SDS mapping.** The SDS split matches coverage: From b7523fcae9c5a48c54a0333d7f6dcccb8a7601ec Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:16:43 +0000 Subject: [PATCH 23/59] =?UTF-8?q?docs:=20rewrite=20=C2=A74.1=20as=20a=20wo?= =?UTF-8?q?rked=20G&L=20example=20and=20=C2=A74.2=E2=80=93=C2=A74.5=20in?= =?UTF-8?q?=20plain=20terms?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 204 ++++++++++-------- 1 file changed, 120 insertions(+), 84 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index bec130fab..ecc6b9d3f 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -111,9 +111,9 @@ A node in the physical data will represent the data or summary instance, so a no | Question | Answer | Section | |---|---|---| | Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | -| What is it based on? | Goldstein & Larson view matching (SIGMOD 2001) | §4.1 | +| What is it based on? | Goldstein & Larson view matching (SIGMOD 2001), explained with an example | §4.1 | | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | -| How does it map to the paper? | computation, aggregate and `GROUP BY` → `definition`; ranges → `selection` | §4.3 | +| What do we take from the paper, and what do we add? | computation, aggregate and `GROUP BY` → `definition`; ranges → `selection`; plus unions and summary families | §4.3 | | Who sets it? | Nobody: it is derived from the sub-DAG | §4.4 | | What uses it? | merge, rollup, slice, reuse, subtract | §4.5 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.6 | @@ -135,149 +135,185 @@ So the schema says *what kind of state* this is, and coverage says *which data i **Why coverage is a sub-DAG, not a few table columns.** A summary state summarizes the result of a whole computation: a KLL over `rate(requests_total[5m])` summarizes rate outputs, and a KLL over a join summarizes join rows. So coverage describes the state by the sub-DAG below it, split into *what is computed* and *which of its rows were taken*. -### 4.1 Design basis: view matching (Goldstein & Larson) +### 4.1 Background: view matching (Goldstein & Larson) -The design follows the view matching algorithm of Goldstein and Larson, which decides when a query can be answered from a materialized select-project-join-group-by (SPJG) view: +Our design is based on this paper: > J. Goldstein and P.-Å. Larson. *Optimizing Queries Using Materialized Views: A Practical, Scalable Solution.* SIGMOD 2001. -**The algorithm.** It splits a view's `WHERE` into three parts: +**The problem it solves.** A materialized view is a query whose result is stored. When a new query arrives, the optimizer wants to answer it from the stored result instead of the base tables. It has to decide two things: does the view contain every row the query needs, and can the answer be computed from the view's output? -- **Column equivalence classes.** An equivalence class is a set of columns that have the same value in every row that satisfies the `WHERE`. They come from column equalities: `a = b AND b = c` gives the class `{a, b, c}`. For example, the join condition `orders.cust_id = customers.id` puts both columns in one class. The algorithm uses the classes to recognize that a query and a view mean the same thing even when they name different but equal columns. -- **A range per column.** The range of a column (more precisely, of an equivalence class) is the interval of values it can have in rows that satisfy the `WHERE`. It comes from comparisons with a constant: `x > 5 AND x <= 10` gives `x ∈ (5, 10]`, and `x = 5` gives `x ∈ [5, 5]`. A column with no such comparison has the range `(−∞, +∞)`. Ranges let the algorithm prove containment: a view with `x ∈ (0, 100]` holds every row a query with `x ∈ (5, 10]` needs. -- **Residual predicates.** A residual predicate is a conjunct of the `WHERE` (one of its `AND`-ed terms) that is neither a column equality nor a comparison of a column with a constant, such as `a + b > 10`, `lower(name) LIKE 'a%'`, or `x = 1 OR y = 2`. The algorithm does not interpret them; it only checks whether the same predicate appears in the query. +**Example.** View `V` is stored; query `Q` arrives: -A view can answer a query when all of these hold: +```sql +-- V: stored +SELECT region, job, day, SUM(bytes) AS s +FROM t WHERE day BETWEEN 1 AND 31 +GROUP BY region, job, day; -1. every residual predicate of the view also appears in the query (§3.1.2, residual subsumption, checked by matching the predicates' text after normalization). Residuals cannot be reasoned about, so the view must not filter out any row the query needs. For example, a view with `WHERE lower(name) LIKE 'a%'` can answer a query only if the query has the same `lower(name) LIKE 'a%'`. The query may have extra residuals; they are applied to the view's output as compensating predicates; -2. the query's ranges lie inside the view's (§3.1.2); -3. the columns needed by compensating predicates are in the view output (§3.3, requirement 2); -4. the query's `GROUP BY` is a subset of the view's, so the query's groups are further aggregations of the view's groups (§3.3, requirement 3). +-- Q: new query +SELECT job, SUM(bytes) +FROM t WHERE day BETWEEN 5 AND 10 AND region = 'us' +GROUP BY job; +``` -**Existing implementation.** The SPJ part is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`). It rejects `Aggregate` and `Join` input plans. +**Step 1: split each `WHERE` into three parts.** Each part is a list of `AND`-ed conditions: -### 4.2 Summary Coverage = Summary definition + selection +| Part | What it is | Comes from | In `V` | In `Q` | +|---|---|---|---|---| +| **Equivalence classes** | sets of columns that are equal in every row | column equalities, e.g. the join condition `orders.cust_id = customers.id` | none | none | +| **Ranges** | for each column, the interval of values it may have | comparisons with a constant: `x > 5`, `x = 5`, `BETWEEN` | `day ∈ [1, 31]` | `day ∈ [5, 10]`, `region ∈ ['us', 'us']` | +| **Residuals** | every other condition; the algorithm does not try to understand them | e.g. `a + b > 10`, `lower(name) LIKE 'a%'`, `x = 1 OR y = 2` | none | none | -A state built by `SummaryAgg` means +**Step 2: does `V` contain every row `Q` needs?** (§3.1.2 of the paper) -```text -state_g = family( input( σ( C ) ) ) for each group value g of G, restricted to G = g +- **Ranges:** each range of `Q` lies inside the same column's range in `V`. `day ∈ [5, 10]` is inside `[1, 31]` ✓. `V` has no range on `region`, so any `region` is in `V` ✓. +- **Residuals:** each residual of `V` also appears in `Q`. Since residuals are not understood, the only safe case is when `Q` has the same condition. +- **Equivalence classes:** each column equality of `V` also holds in `Q`. + +**Step 3: can the answer be computed from `V`'s output?** (§3.3) + +- **Compensating filter:** where `Q` is narrower than `V`, the extra condition is applied to `V`'s rows. So its columns must be in `V`'s output: `day` and `region` are ✓. +- **Regrouping:** `Q`'s `GROUP BY` must be a subset of `V`'s. `{job}` ⊆ `{region, job, day}` ✓, so each group of `Q` is the sum of some groups of `V`. + +**Result:** + +```sql +SELECT job, SUM(s) FROM V +WHERE day BETWEEN 5 AND 10 AND region = 'us' +GROUP BY job; ``` -- `C` is the child sub-DAG with the selection removed. Its output rows are the contributions. -- `σ` is the selection: which output rows of `C` went into the state. +**Limits that matter for us:** + +- It answers a query from **one** view. Combining several views (a union) is left out (§3.1). +- It supports only `SUM` and `COUNT`, whose groups can be added up again. + +**Existing implementation.** The `WHERE` split is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`). It rejects plans that contain an `Aggregate` or a `Join`. -Coverage stores exactly these two things: +### 4.2 Summary Coverage = Summary definition + selection + +A summary state is a stored aggregation, like `V` above, whose aggregate is a sketch. So we describe it the way the paper describes a view, in two parts: -| Part | Contents | Meaning | +| Part | Question it answers | What it is | |---|---|---| -| **`definition`** | the `SummaryAgg` node itself, with the selection removed from its child sub-DAG; carries `C`, `input`, `family` and `G` | what the state *means* | -| **`selection`** | a union of boxes over the output columns of `C` | *which rows* the state took | +| **`definition`** | *What* is computed? | the `SummaryAgg` node with its row filters taken out: the sub-DAG below it, the summary family and parameters, its input column, and its `GROUP BY` | +| **`selection`** | *Which rows* went in? | the row filters that were taken out, as simple conditions on columns | -**Why this is enough.** If two states have the same `definition`, their contributions come from the same rows of the same computation, whatever `C` contains (join, union, `rate`, dedup). Disjoint selections then cannot share a row, so no observation is counted twice. No per-operator occurrence rule is needed. +**Example:** -**Examples** of what ends up where: +```text +state = KLL(latency) by[job] over Filter(region = 'us' AND latency < 100, Scan t) + +definition: KLL(latency) by[job] over Scan t +selection: region ∈ {us}, latency ∈ (−∞, 100) +``` + +**Why this is enough.** Two states with the same `definition` come from the same computation. If their selections do not overlap, no row is in both, so merging them counts every row once. This holds whatever the computation contains (joins, unions, `rate`, dedup), so we need no special rule per operator. + +**More examples** of what goes where: | Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | |---|---|---| | `Filter(region = 'us', Scan t)` | `Scan t` | `region ∈ {us}` | | `Filter(latency < 100, Scan t)` | `Scan t` | `latency ∈ (−∞, 100)` | -| `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | time `(−3m, −2m]` relative to evaluation | -| `Filter(rate > 0, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))`, `rate > 0` as residual | — | +| `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | the last 3 to 2 minutes before evaluation, `(−3m, −2m]` | | `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | +| `Filter(rate > 0, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `rate > 0` | nothing | -In the last two rows the `TimeRange(5m)` stays in `definition`: it sits below `rate` and changes the rate values, so it is not a selection of output rows. Reading those 5 minutes of source data is the runtime's job, not coverage (§4.7). +In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `rate > 0` stays too, because it filters on a computed value (§4.4). -### 4.3 Mapping to Goldstein & Larson +Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` is the sub-DAG without its row filters, `σ` is the selection, and `G` the grouping. -A summary state is an aggregation view whose aggregate is a summary family: +### 4.3 What we take from Goldstein & Larson, and what we add | Goldstein & Larson | Summary coverage | |---|---| -| SPJ part: tables, joins, residual predicates | the computation `C` below the `SummaryAgg` (§4.2), part of `definition` | -| aggregate function and its argument | `family` and `input` (`SummaryUpdate`) of the `SummaryAgg`, part of `definition` | -| `GROUP BY` | the `SummaryAgg` reduction `G`, part of `definition` | -| ranges per column | `selection` (§4.4) | -| compensating predicate on view output | slice on a column of `G` only (§4.5) | -| query `GROUP BY` ⊆ view `GROUP BY` | rollup (§4.5) | +| tables, joins and residuals of the view | the sub-DAG `C` below the `SummaryAgg`, in `definition` | +| the aggregate and its argument | the summary family and its input, in `definition` | +| `GROUP BY` | the `SummaryAgg` grouping `G`, in `definition` | +| ranges | `selection` (§4.4) | +| compensating filter on the view's output | slicing: allowed only on a column of `G` (§4.5) | +| regrouping to a smaller `GROUP BY` | rollup (§4.5) | -**What this design adds beyond the paper**: +**What we add:** -- **Unions of states.** The paper considers single-view substitutes and notes that its requirement 1, that the view contains all rows the query needs, "is not required if substitutes containing unions of views are considered" (§3.1). `SummaryMerge` is exactly such a union, so it needs a disjointness check the paper does not have. -- **Summary families.** The paper allows `SUM` and `COUNT_BIG` only. Here each family declares how the selections of its inputs may relate (§4.5). -- **Value sets and hash partitions** next to ranges, and **evaluation-relative time** (§4.4). +- **Unions of states.** The paper uses one view at a time. `SummaryMerge` combines several states, so we must also check that their selections do not overlap. +- **Summary families.** The paper only re-adds `SUM` and `COUNT`. Each summary family says how its inputs may overlap (§4.5). +- **More kinds of conditions:** value sets (`IN`, `NOT IN`), hash partitions, and time relative to the evaluation time (§4.4). ### 4.4 Deriving the selection -Coverage is derived from the node, never declared. Walking down from the `SummaryAgg` (its own `filter` included), a predicate conjunct goes into `selection` when both rules hold. Otherwise it stays in `definition` as a residual, as in Goldstein & Larson. +Nobody declares coverage: the planner computes it from the sub-DAG. It starts at the `SummaryAgg`, walks down, and looks at each filter condition on the way (split at `AND`). A condition moves into `selection` only when both rules below hold. Otherwise it stays in `definition`, like a residual in the paper. -**Rule 1: it can be lifted to the `SummaryAgg`.** Lifting is the inverse of DataFusion's `PushDownFilter` (`datafusion-optimizer`, `push_down_filter.rs`): +**Rule 1: the condition can move up to the `SummaryAgg` without changing its meaning.** This is the reverse of filter pushdown (DataFusion's `PushDownFilter`). -| Operator on the path | The predicate passes when | +| Operator between the condition and the `SummaryAgg` | The condition can pass when | |---|---| | `Filter`, `Scan.predicates`, `SummaryAgg.filter` | always | | range `TimeRange`, `TimeShift` without `@` | always | -| `Project` | the column is a direct column item (renaming keeps its identity) | -| `Aggregate` (later) | every column it uses is a group column | -| window function, per-series temporal function such as `rate` (later) | every column it uses is a partition column (a series label) | +| `Project` | the column is passed through as is (a rename is fine) | +| `Aggregate` (later) | it uses only group columns | +| window function, or a per-series function such as `rate` (later) | it uses only partition columns (series labels) | | anything else | never | -**Rule 2: it is a box constraint on one column.** +**Rule 2: the condition is a simple condition on one column.** -| Constraint | From | DataFusion analogue | -|---|---|---| -| value set: `In` / `NotIn` literals | `=`, `!=`, `IN`, `NOT IN`, `OR` of equalities | `LiteralGuarantee` | -| interval: lower and upper `std::ops::Bound` (`Included`, `Excluded`, `Unbounded`) | `<`, `<=`, `>`, `>=` | `Interval` | -| hash partition (later): `hash(columns) mod n = k` | partitioned producers | `Partitioning::Hash` | +| Kind | Written as | Example | DataFusion analogue | +|---|---|---|---| +| value set | `=`, `!=`, `IN`, `NOT IN`, `OR` of `=` on the same column | `region IN ('us', 'eu')` | `LiteralGuarantee` | +| interval | `<`, `<=`, `>`, `>=` | `latency < 100` | `Interval` | +| hash partition (later) | `hash(columns) mod n = k` | `hash(job) mod 4 = 1` | `Partitioning::Hash` | -Column equalities (`a = b`) are residuals: there are no column equivalence classes. +A column equality such as `a = b` is not a simple condition, so it stays in `definition`. -**Column identity.** +**Which column a condition is on.** -- Columns are identified by lineage `(table, name)`, the identity `ColumnRef::Qualified` uses, not by `Field.name`. So `shipping.region` and `billing.region` stay different columns. -- A direct alias keeps the identity of the column it renames. -- A column whose `(table, name)` is not unique in the output (two items aliased `k`) cannot be named, so its conjuncts stay residual. -- Value sets compare literals by type: `1` and `1.0` are never proven different. +- A column is named by its source table and name, `(table, name)`. So `shipping.region` and `billing.region` are different columns. +- A rename keeps the original name. +- If two output columns have the same `(table, name)`, the condition cannot tell them apart and stays in `definition`. +- Values of different types are never treated as different: `1` and `1.0` might be equal. -**Time** is a selection like any other: +**Time.** There are two kinds: -| Kind | How it is selected | Example | +| Kind | Where it comes from | Example | |---|---|---| -| **Absolute** | an interval on the timestamp column (the schema's `time_index`) | `ts >= t0 AND ts < t1` → `(Included(t0), Excluded(t1))` on `ts` | -| **Relative** | its own field `relative_time`, because it is not a column value: one range `TimeRange(w)` over a `TimeShift(s)` | `(Excluded(−(s+w)), Included(−s))` relative to evaluation | +| **Absolute** | an interval on the timestamp column | `ts >= t0 AND ts < t1` → `ts ∈ [t0, t1)` | +| **Relative** to the evaluation time | a range `TimeRange(w)` over a `TimeShift(s)` | `TimeRange(1m)` over `TimeShift(2m)` → `(−3m, −2m]` | -- PromQL ranges are left-open, matching the executor (`series_window.rs`). -- The IR has no timestamp literal yet, so absolute SQL time filters stay residual until it does. -- Stage 2 tumbling panes (`window_composition.rs` in #601) get their time this way, so it is derived rather than declared. -- An instant `TimeRange` picks the latest sample per series, which is not a selection of rows, so it stays in `definition`. -- Relative time and a timestamp-column interval are different dimensions, so they are never compared: two states restricted only by different kinds of time are treated as possibly overlapping. -- Binding a relative pane to absolute timestamps for one evaluation (evaluation time plus the pane layout's phase) is a runtime coordinate, not coverage. +- PromQL windows exclude their start, so relative windows are open on the left. +- The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. +- Stage 2 builds its tumbling panes from `TimeRange` and `TimeShift`, so pane times are derived, not declared (#601). +- An instant `TimeRange` takes the latest sample of each series. That does not pick rows by time, so it stays in `definition`. +- Absolute and relative time are never compared. Two states that differ only in the kind of time are treated as possibly overlapping. ### 4.5 Operations -| Operation | Example | Valid when | +Coverage tells the planner which states can be combined, and what the result covers. + +| Operation | Example | Allowed when | |---|---|---| -| merge (`SummaryMerge`, same `G`) | `[0,1m)` ⊕ `[1m,2m)`; `region='us'` ⊕ `region='eu'` | all `definition`s equal; selections related as the family requires (below) | -| rollup (`SummaryMerge` with `group_by: G'`, later) | `by[region, job]` → `by[job]` | `G'` ⊆ `G` and the family merges. Groups of one state are disjoint because a row has one value per group column, so no selection check is needed | -| slice | `by[region, job]` state answering `region = 'us' … by[job]` | the restricted columns are all in `G`. A sketch cannot be filtered, so a restriction on any other column is invalid | -| reuse for a query | a stored state answers a query | Goldstein & Larson containment: same `definition`, query selection inside the state's, any compensating restriction is a slice | -| subtract (`SummarySubtract`, reserved) | `[0,10) − [0,5)` | same `definition`; the right selection is contained in the left | +| **merge** (`SummaryMerge`) | minute 0–1 + minute 1–2; `region='us'` + `region='eu'` | all `definition`s are equal, and the selections relate as the family requires (below) | +| **rollup** (`SummaryMerge` with `group_by`, later) | `by[region, job]` → `by[job]` | the new grouping is a subset of the old one. Different groups never share a row, so no overlap check is needed | +| **slice** | read `region = 'us'` from a `by[region, job]` state | the condition is on a grouping column. A sketch cannot be filtered on any other column | +| **reuse** for a query | a stored state answers a new query | same `definition`, and the query's selection lies inside the state's (as in the paper) | +| **subtract** (`SummarySubtract`, reserved) | `[0, 10) − [0, 5)` | same `definition`, and the right selection lies inside the left | -**Merge and rollup are one operator.** `SummaryMerge { children, group_by }`: one child with a coarser `group_by` is a rollup, and `group_by` equal to the children's is a plain merge. Its coverage is the children's `definition` with `G'` and the union of their selections; adjacent intervals are joined, gaps stay as separate boxes. +**Merge and rollup are one operator,** `SummaryMerge { children, group_by }`. With the children's own grouping it is a plain merge; with a smaller one it is a rollup. The result's `definition` is the children's, with the new grouping; its `selection` is the union of theirs. Adjacent ranges join into one (minute 0–1 + minute 1–2 = minute 0–2); gaps stay as separate pieces. -**How selections must relate** is declared by the family, next to whether it merges (`FieldDataType::family_merges` and `merge_relation`, added in #592): +**How inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): -| Relation | Families | Why | +| Rule | Families | Why | |---|---|---| -| **disjoint** | counting families: KLL, Count-Min, exact `Sum`/`Count` | an overlapping row would be counted twice | -| **overlap allowed** | idempotent families: HLL, exact `Min`/`Max`, distinct sets | adding a row twice does not change the state | -| **contained** | subtraction | the right state must be part of the left | +| **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | a row in both inputs would be counted twice | +| **may overlap** | HLL, exact `Min`/`Max`, distinct sets | adding the same row twice does not change the result | +| **right inside left** | subtraction | you can only remove what is there | **When two `definition`s are equal.** -- They must be structurally equal, ignoring planning metadata: `timing`, `guarantee` and `coverage_cache`. So a state built at ingestion time and one built at query time can merge. -- `SummaryUpdate.weight_domain` is compared: it is derived from `C` and `input`, so it differs only if a derivation is wrong. -- States over different sources have different `definition`s and do not merge. To combine tables, put `UNION ALL` with a marker column below one `SummaryAgg`; the marker is then an ordinary column for `selection` or `G`. +- They must have the same structure. Planning details are ignored: `timing`, `guarantee` and `coverage_cache`. So a state built at ingestion time can merge with one built at query time. +- `SummaryUpdate.weight_domain` is compared too. It is computed from the rest, so it differs only if something is wrong. +- States over different tables never merge. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. ### 4.6 Interface From 3e3c9ec82a5c369988bcf4e8d7d634be5ee3a1e7 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:27:28 +0000 Subject: [PATCH 24/59] =?UTF-8?q?docs:=20worked=20examples=20for=20=C2=A74?= =?UTF-8?q?.4=20and=20=C2=A74.5;=20drop=20the=20=C2=A75.8=20recap=20table?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 154 +++++++++++------- 1 file changed, 97 insertions(+), 57 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index ecc6b9d3f..8d5b409a9 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -219,9 +219,9 @@ selection: region ∈ {us}, latency ∈ (−∞, 100) | `Filter(latency < 100, Scan t)` | `Scan t` | `latency ∈ (−∞, 100)` | | `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | the last 3 to 2 minutes before evaluation, `(−3m, −2m]` | | `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | -| `Filter(rate > 0, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `rate > 0` | nothing | +| `Filter(value * 2 > 10, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `value * 2 > 10` | nothing | -In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `rate > 0` stays too, because it filters on a computed value (§4.4). +In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column (§4.4). Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` is the sub-DAG without its row filters, `σ` is the selection, and `G` the grouping. @@ -244,76 +244,131 @@ Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` ### 4.4 Deriving the selection -Nobody declares coverage: the planner computes it from the sub-DAG. It starts at the `SummaryAgg`, walks down, and looks at each filter condition on the way (split at `AND`). A condition moves into `selection` only when both rules below hold. Otherwise it stays in `definition`, like a residual in the paper. +Nobody declares coverage: the planner computes it from the sub-DAG. It starts at the `SummaryAgg`, walks down, and looks at each filter condition on the way (split at `AND`). A condition moves into `selection` only when both rules hold: -**Rule 1: the condition can move up to the `SummaryAgg` without changing its meaning.** This is the reverse of filter pushdown (DataFusion's `PushDownFilter`). +- **Rule 1: it can move up to the `SummaryAgg` without changing its meaning.** +- **Rule 2: it is a simple condition on one column.** -| Operator between the condition and the `SummaryAgg` | The condition can pass when | -|---|---| -| `Filter`, `Scan.predicates`, `SummaryAgg.filter` | always | -| range `TimeRange`, `TimeShift` without `@` | always | -| `Project` | the column is passed through as is (a rename is fine) | -| `Aggregate` (later) | it uses only group columns | -| window function, or a per-series function such as `rate` (later) | it uses only partition columns (series labels) | -| anything else | never | +Otherwise it stays in `definition`, like a residual in the paper. + +**Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: + +```text +SummaryAgg KLL(value) by[job] filter: value < 100 +└─ Project [job, region AS r, value] + └─ Filter region = 'us' AND value * 2 > 10 + └─ TimeRange 1m (range) + └─ TimeShift 2m + └─ Scan m +``` + +| Condition | Found at | Rule 1: moves up? | Rule 2: simple? | Result | +|---|---|---|---|---| +| `value < 100` | `SummaryAgg.filter` | yes, it is already there | yes, an interval | `selection`: `value ∈ (−∞, 100)` | +| `region = 'us'` | `Filter` | yes: `Project` passes `region` through (renamed `r`) | yes, a value set | `selection`: `m.region ∈ {us}` | +| `value * 2 > 10` | `Filter` | yes | **no**: it is on an expression, not a column | stays in `definition` | +| 1 minute, shifted by 2 | `TimeRange` + `TimeShift` | yes | yes, relative time | `selection`: `(−3m, −2m]` | -**Rule 2: the condition is a simple condition on one column.** +Result: + +```text +definition: KLL(value) by[job] over Project[job, region AS r, value] over Filter(value * 2 > 10) over Scan m +selection: value ∈ (−∞, 100), m.region ∈ {us}, time (−3m, −2m] +``` + +**Rule 1 in detail.** This is the reverse of filter pushdown (DataFusion's `PushDownFilter`). A condition can move up past: + +| Operator | Moves up? | Example | +|---|---|---| +| `Filter`, `Scan.predicates`, `SummaryAgg.filter` | always | above | +| range `TimeRange`, `TimeShift` without `@` | always | above | +| `Project` | only for a column passed through as is (a rename is fine) | `region AS r` ✓; `value * 2 AS v2` ✗ | +| `Aggregate` (later) | only on group columns | below `SUM(value) by[job]`: `job = 'api'` ✓, `value > 5` ✗ | +| window function, or per-series function such as `rate` (later) | only on series labels | below `rate(...)`: `job = 'api'` ✓, `value > 5` ✗ (it would filter raw samples, which changes the rate) | +| anything else | never | | -| Kind | Written as | Example | DataFusion analogue | +A condition *above* `rate` is different: it filters the rate outputs, which are exactly the rows the `SummaryAgg` sees. `Filter(job = 'api', rate(...))` gives `job ∈ {api}`, and `Filter(value > 0, rate(...))` gives `value ∈ (0, ∞)` on the rate values. The `TimeRange(5m)` below `rate` stays in `definition` either way. + +**Rule 2 in detail.** A simple condition is one of: + +| Kind | Written as | Example | Becomes | |---|---|---|---| -| value set | `=`, `!=`, `IN`, `NOT IN`, `OR` of `=` on the same column | `region IN ('us', 'eu')` | `LiteralGuarantee` | -| interval | `<`, `<=`, `>`, `>=` | `latency < 100` | `Interval` | -| hash partition (later) | `hash(columns) mod n = k` | `hash(job) mod 4 = 1` | `Partitioning::Hash` | +| value set | `=`, `!=`, `IN`, `NOT IN`, `OR` of `=` on one column | `region IN ('us', 'eu')` | `region ∈ {us, eu}` | +| | | `region != 'test'` | `region ∉ {test}` | +| interval | `<`, `<=`, `>`, `>=` | `value >= 10 AND value < 100` | `value ∈ [10, 100)` | +| hash partition (later) | `hash(columns) mod n = k` | `hash(job) mod 4 = 1` | partition 1 of 4 | -A column equality such as `a = b` is not a simple condition, so it stays in `definition`. +Not simple, so they stay in `definition`: `value * 2 > 10` (expression), `a = b` (two columns), `region = 'us' OR job = 'api'` (two columns), `name LIKE 'web%'` (pattern). **Which column a condition is on.** -- A column is named by its source table and name, `(table, name)`. So `shipping.region` and `billing.region` are different columns. -- A rename keeps the original name. -- If two output columns have the same `(table, name)`, the condition cannot tell them apart and stays in `definition`. -- Values of different types are never treated as different: `1` and `1.0` might be equal. +- A column is named by its source table and name, `(table, name)`: in a join, `shipping.region = 'us'` and `billing.region = 'us'` are different conditions. +- A rename keeps the original name: `region AS r` is still `m.region`. +- If two output columns have the same `(table, name)` (for example `Project [a AS k, b AS k]`), a condition on `k` cannot tell them apart and stays in `definition`. +- Values of different types are never treated as different: `1` and `1.0` might be equal, so `x = 1` and `x = 1.0` are treated as possibly overlapping. **Time.** There are two kinds: -| Kind | Where it comes from | Example | +| Kind | Comes from | Example | |---|---|---| +| **Relative** to the evaluation time | a range `TimeRange(w)` over a `TimeShift(s)` → `(−(s+w), −s]` | `TimeRange(1m)` alone → `(−1m, 0]`; over `TimeShift(1m)` → `(−2m, −1m]` | | **Absolute** | an interval on the timestamp column | `ts >= t0 AND ts < t1` → `ts ∈ [t0, t1)` | -| **Relative** to the evaluation time | a range `TimeRange(w)` over a `TimeShift(s)` | `TimeRange(1m)` over `TimeShift(2m)` → `(−3m, −2m]` | - PromQL windows exclude their start, so relative windows are open on the left. +- Stage 2 builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). +- An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. -- Stage 2 builds its tumbling panes from `TimeRange` and `TimeShift`, so pane times are derived, not declared (#601). -- An instant `TimeRange` takes the latest sample of each series. That does not pick rows by time, so it stays in `definition`. -- Absolute and relative time are never compared. Two states that differ only in the kind of time are treated as possibly overlapping. +- Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. ### 4.5 Operations -Coverage tells the planner which states can be combined, and what the result covers. +Coverage tells the planner which states can be combined, and what the result covers. The examples below use these states. All are `KLL(value) by[job] over Scan m` unless noted: -| Operation | Example | Allowed when | -|---|---|---| -| **merge** (`SummaryMerge`) | minute 0–1 + minute 1–2; `region='us'` + `region='eu'` | all `definition`s are equal, and the selections relate as the family requires (below) | -| **rollup** (`SummaryMerge` with `group_by`, later) | `by[region, job]` → `by[job]` | the new grouping is a subset of the old one. Different groups never share a row, so no overlap check is needed | -| **slice** | read `region = 'us'` from a `by[region, job]` state | the condition is on a grouping column. A sketch cannot be filtered on any other column | -| **reuse** for a query | a stored state answers a new query | same `definition`, and the query's selection lies inside the state's (as in the paper) | -| **subtract** (`SummarySubtract`, reserved) | `[0, 10) − [0, 5)` | same `definition`, and the right selection lies inside the left | +| State | Selection | +|---|---| +| `A` | time `(−1m, 0]` | +| `B` | time `(−2m, −1m]` | +| `C` | time `(−90s, −30s]` | +| `D` | time `(−1m, 0]`, but the `definition` has the residual `value * 2 > 10` | +| `E` | time `(−1m, 0]`, `region ∈ {us}` | +| `F` | time `(−1m, 0]`, `region ∈ {eu}` | + +**Merge** (`SummaryMerge`): combine states into one. Allowed when all `definition`s are equal and the selections relate as the family requires. -**Merge and rollup are one operator,** `SummaryMerge { children, group_by }`. With the children's own grouping it is a plain merge; with a smaller one it is a rollup. The result's `definition` is the children's, with the new grouping; its `selection` is the union of theirs. Adjacent ranges join into one (minute 0–1 + minute 1–2 = minute 0–2); gaps stay as separate pieces. +| Merge | Allowed? | Why | Result's selection | +|---|---|---|---| +| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (adjacent ranges join) | +| `E + F` | ✓ | same definition, `us` and `eu` do not overlap | `(−1m, 0]`, `region ∈ {us, eu}` | +| `A + C` | ✗ | `(−60s, −30s]` is in both: those rows would be counted twice | | +| `A + A` | ✗ | every row is in both | | +| `A + D` | ✗ | different definitions: `D` only has rows with `value * 2 > 10` | | +| `A + B'` where `B'` is KLL with `k = 400` | ✗ | different definitions (parameters) | | +| `(A + B) + B''` where `B''` covers `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | **How inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): -| Rule | Families | Why | +| Rule | Families | Example | |---|---|---| -| **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | a row in both inputs would be counted twice | -| **may overlap** | HLL, exact `Min`/`Max`, distinct sets | adding the same row twice does not change the result | -| **right inside left** | subtraction | you can only remove what is there | +| **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | +| **may overlap** | HLL, exact `Min`/`Max`, distinct sets | HLL over `A`'s and `C`'s selections ✓: a value seen twice is still one distinct value; the result covers `(−90s, 0]` | +| **right inside left** | subtraction | see subtract below | + +**Rollup** (`SummaryMerge` with `group_by`, later): make the grouping coarser. A `by[region, job]` state rolls up to `by[job]`: the state for `job = api` is the merge of `(us, api)`, `(eu, api)`, …. No overlap check is needed, because a row has one `region` and so is in only one group. The selection is unchanged. + +**Slice**: read only some groups. From a `by[region, job]` state, a query for `region = 'us'` by job reads the groups with `region = us` ✓. A query for `value < 50` ✗: `value` is not a grouping column, and a sketch cannot be filtered after it is built. + +**Reuse** for a query: a stored state answers a query when the `definition`s are equal and the query's rows are all in the state, with any difference covered by a slice. A stored `by[region, job]` state over `(−5m, 0]`: + +- p99 by job over the last 5 minutes for `region = 'us'`: ✓ (slice on `region`). +- p99 by job over the last 1 minute: ✗. The state also holds minutes 2–5, and time is not a grouping column, so they cannot be taken out. + +**Subtract** (`SummarySubtract`, reserved): remove one state from another, for families that allow it (e.g. exact `Sum`/`Count`, Count-Min). A sum over `(−10m, 0]` minus a sum over `(−10m, −5m]` gives `(−5m, 0]`. Allowed when the `definition`s are equal and the right selection is inside the left. **When two `definition`s are equal.** -- They must have the same structure. Planning details are ignored: `timing`, `guarantee` and `coverage_cache`. So a state built at ingestion time can merge with one built at query time. +- They must have the same structure. Planning details are ignored: `timing`, `guarantee` and `coverage_cache`. So a pane built at ingestion time can merge with one built at query time. - `SummaryUpdate.weight_domain` is compared too. It is computed from the rest, so it differs only if something is wrong. -- States over different tables never merge. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. +- States over different tables never merge: a KLL over `m1` and one over `m2` have different definitions. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. ### 4.6 Interface @@ -520,21 +575,6 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | | `Extension` | `child, name` | deployment-named state operator | -### 5.8 Summary - -| Operator | Input kind | Output kind | Output carries state | Coverage on output | Status | -|---|---|---|---|---|---| -| `SummaryAgg` | value (not `State`) | `State` | yes (one `family` field) | derived: itself minus selection, plus selection | implemented | -| `SummaryEstimate` | `State` (one `Sketch` field) | source's value kind | no | none | implemented | -| `FinalizeExactAccumulator` | `State` (`ExactAggregate`) | source's value kind | no | none | implemented | -| `MaintainPopulation` | `Relation` (table) / `InstantVector` (series) | `State` | yes (by kind; fields plain) | none | implemented | -| `EvaluatePopulation` | `State` from `MaintainPopulation` | source's value kind | no | none | implemented | -| `SummaryMerge` | `State` × N | `State` | yes | derived: shared definition with `group_by`, union of selections | implemented (#560); coverage check in #646; `group_by` planned | -| `SummarySubtract` | `State` × 2 | `State` | yes | derived: left selection minus right (planned) | reserved | -| `SummaryDelete` | `State` | `State` | yes | — | reserved | -| `SummaryJoin` | `State` × 2 | `State` | yes | — | reserved | -| `Extension` | any | `State` | yes | — | reserved | - ## 6. Key code interfaces `OperatorNode`, `OperatorResultKind` and coverage are in §4. Bodies and serde/derive attributes are elided below. From bfa0111ad8a7b8e3f8dbc60384988c8d7e2a497f Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 18:28:39 +0000 Subject: [PATCH 25/59] =?UTF-8?q?docs:=20draw=20the=20=C2=A75=20examples?= =?UTF-8?q?=20as=20diagrams=20and=20timelines?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 208 ++++++++++++++---- 1 file changed, 160 insertions(+), 48 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 8d5b409a9..261563156 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -435,39 +435,55 @@ Coverage only says what a state means and which rows it took. The deployment and Given that these information requirements are introduced by summary operators to work correctly semantically, we show the examples of how the defined OperatorNode, schema, and physical data information work with each kind of summary operators. -Notation: an edge is written `──Kind(field Type, …)──▶`. Schemas are the ones `output_schema()` derives. Planning may rename fields through `OperatorNode::with_schema`, but types, nullability, `time_index`, `unique_keys` and `closed` must match the derivation. All examples use a table source, so values are `Relation`; with a `TimeSeries` source the value side is `InstantVector`. +How to read the diagrams: data flows from top to bottom. Each edge is labelled with the schema it carries, written `Kind: field Type, …`. -### 5.1 `SummaryAgg`: values → state +| Color | Meaning | +|---|---| +| blue box | operator whose output is a value (`Relation`) | +| yellow box | operator whose output is summary state (`State`) | +| gray dashed box | the node's `coverage()` | -Scenario: p99 latency by job, from KLL(k=200), over one minute of table `t`, US rows only. +Schemas are the ones `output_schema()` derives. Planning may rename fields through `OperatorNode::with_schema`, but types, nullability, `time_index`, `unique_keys` and `closed` must match the derivation. All examples use a table source, so values are `Relation`; with a `TimeSeries` source the value side is `InstantVector`. -```text -Scan(t: job Utf8, region Utf8, ts Timestamp [time_index], latency Float64) - ──Relation(job Utf8, region Utf8, ts Timestamp, latency Float64)──▶ -Filter(region = 'us' AND ts >= 0 AND ts < 60_000) - ──Relation(job Utf8, region Utf8, ts Timestamp, latency Float64)──▶ -SummaryAgg(family = Sketch(KLL{k=200}, PerSubpopulationInstance), - input = SummaryUpdate::column(Named("latency")), reduction = by[job], - grouping = PerSubpopulationInstance, filter = None) - ──State(job Utf8, state Sketch(KLL{k=200}, PerSubpopulationInstance))──▶ - coverage() = { definition: this SummaryAgg over Scan(t) (the Filter removed), - selection: [{ columns: { t.region: In{'us'} }, - time: Absolute [Included(0), Excluded(60_000)) }] } +### 5.1 `SummaryAgg`: values → state + +Scenario: p99 latency by job, from KLL(k=200), over table `t`, US rows with latency under 10 s only. + +```mermaid +flowchart TB + scan["Scan t"]:::value + filter["Filter
region = 'us' AND latency #lt; 10000"]:::value + agg["SummaryAgg
family: Sketch KLL k=200
input: column latency
reduction: by job"]:::state + out(["next operator"]) + cov["coverage()
definition: this SummaryAgg over Scan t (Filter removed)
selection: t.region ∈ {us}, t.latency ∈ (−∞, 10000)"]:::cov + scan -->|"Relation: job Utf8, region Utf8, ts Timestamp, latency Float64"| filter + filter -->|"Relation: same as above"| agg + agg -->|"State: job Utf8, state Sketch(KLL k=200)"| out + agg -.- cov + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 ``` - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. - Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **always derived**, never declared (§4.4). Both conjuncts of the `Filter` lift into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. +- Coverage: **always derived**, never declared (§4.4). Both conditions of the `Filter` move into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. - Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. ### 5.2 `SummaryEstimate`: sketch state → value Scenario: read p99 from the state in 5.1. -```text -──State(job Utf8, state Sketch(KLL{k=200}))──▶ -SummaryEstimate(query = SketchStatistic::Quantile { q: 0.99 }) - ──Relation(job Utf8, quantile Float64)──▶ (planner may rename to p99) +```mermaid +flowchart TB + agg["SummaryAgg from 5.1"]:::state + est["SummaryEstimate
query: Quantile q = 0.99"]:::value + out(["next operator"]) + agg -->|"State: job Utf8, state Sketch(KLL k=200)"| est + est -->|"Relation: job Utf8, quantile Float64 (planner may rename to p99)"| out + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 ``` - Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`). Keys and metadata pass through. A top-k readout is the exception: it returns the selected rows, one per ranked item, with the partition keys, the item identity columns, and a `value` Float64 score (#579). This is the same row shape as an exact Sort → Limit top-k, so the plans for one query share a root schema. @@ -480,13 +496,20 @@ SummaryEstimate(query = SketchStatistic::Quantile { q: 0.99 }) Scenario: total bytes by host with an exact Sum accumulator. -```text -Scan(t: host Utf8, bytes Float64) - ──Relation(host Utf8, bytes Float64)──▶ -SummaryAgg(family = ExactAggregate(Sum, Sum), input = column(Named("bytes")), reduction = by[host]) - ──State(host Utf8, state ExactAggregate(Sum, Sum))──▶ coverage: derived -FinalizeExactAccumulator - ──Relation(host Utf8, state Float64)──▶ +```mermaid +flowchart TB + scan["Scan t"]:::value + agg["SummaryAgg
family: ExactAggregate Sum
input: column bytes
reduction: by host"]:::state + fin["FinalizeExactAccumulator"]:::value + out(["next operator"]) + cov["coverage()
definition: this SummaryAgg over Scan t
selection: everything"]:::cov + scan -->|"Relation: host Utf8, bytes Float64"| agg + agg -->|"State: host Utf8, state ExactAggregate(Sum)"| fin + fin -->|"Relation: host Utf8, state Float64"| out + agg -.- cov + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 ``` - Output schema: each `ExactAggregate` field keeps its name (`state`) and takes the type and nullability the equivalent `NonASAPOp::Aggregate` would give: Sum/Min/Max follow the input column, Count is Int64, and Rate/IRate/Increase are Float64. If the child is not a `SummaryAgg` directly, Count falls back to Int64 and the others to Float64. `unique_keys`, `closed` and `time_index` are preserved (`schema_rebuilding.rs`). @@ -498,13 +521,16 @@ FinalizeExactAccumulator Scenario: keep the full latency population per job, so that p99 and top-10 can be evaluated later. -```text -Scan(t: job Utf8, latency Float64) [closed schema] - ──Relation(job Utf8, latency Float64)──▶ -MaintainPopulation(population = MaintainedPopulation { - input: PopulationInput::Rows { input: , value_column: 1, grouping: by[job] }, - max_k: 10, quantiles: true }) - ──State(job Utf8, latency Float64)──▶ +```mermaid +flowchart TB + scan["Scan t (closed schema)"]:::value + mp["MaintainPopulation
input: Rows of that Scan, value column latency, by job
max_k: 10, quantiles: true"]:::state + out(["next operator"]) + scan -->|"Relation: job Utf8, latency Float64"| mp + mp -->|"State: job Utf8, latency Float64 (all plain)"| out + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 ``` - Output schema: identical to the child's, all plain. Only `result_kind = State` marks it as maintained state. @@ -516,10 +542,16 @@ MaintainPopulation(population = MaintainedPopulation { Scenario: p99 by job from the population in 5.4. -```text -──State(job Utf8, latency Float64) [from MaintainPopulation]──▶ -EvaluatePopulation(evaluation = PopulationStatistic::Quantile { q: 0.99 }) - ──Relation(job Utf8, quantile_0_99 Float64)──▶ +```mermaid +flowchart TB + mp["MaintainPopulation from 5.4"]:::state + ev["EvaluatePopulation
evaluation: Quantile q = 0.99"]:::value + out(["next operator"]) + mp -->|"State: job Utf8, latency Float64"| ev + ev -->|"Relation: job Utf8, quantile_0_99 Float64"| out + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 ``` - Output schema: the schema of `Aggregate(by grouping, measure)` over the maintained source. Quantile gives `quantile_` Float64, Sum gives `sum` (value type), Count gives `count` Int64 and Average gives `avg` Float64; `unique_keys = [[0]]`, `closed`. `TopK { k }` instead returns the source schema unchanged (the selected rows). @@ -539,21 +571,92 @@ Current state: SummaryMerge { children: Vec, group_by: Reduction } ``` -Scenario A, time panes: two one-minute KLL panes of PromQL `quantile_over_time(0.99, m[2m])` merged into the two-minute state. Each pane reads `TimeRange(1m)` over `TimeShift(s)` over the scan, as Stage 2 builds them. +**Scenario A, time panes.** Two one-minute KLL panes of PromQL `quantile_over_time(0.99, m[2m])` merge into the two-minute state. Each pane reads `TimeRange(1m)` over `TimeShift(s)` over the scan, as Stage 2 builds them. Both panes share one `Scan`. + +```mermaid +flowchart TB + scan["Scan m"]:::value + s0["TimeShift 0"]:::value + s1["TimeShift 1m"]:::value + r0["TimeRange 1m"]:::value + r1["TimeRange 1m"]:::value + p0["pane 0: SummaryAgg
KLL k=200, by nothing"]:::state + p1["pane 1: SummaryAgg
KLL k=200, by nothing"]:::state + m["SummaryMerge
group_by: nothing"]:::state + out(["next operator"]) + c0["selection: time (−1m, 0]"]:::cov + c1["selection: time (−2m, −1m]"]:::cov + cm["selection: time (−2m, 0]
definition: same as both panes"]:::cov + scan --> s0 --> r0 --> p0 + scan --> s1 --> r1 --> p1 + p0 -->|"State: state Sketch(KLL k=200)"| m + p1 -->|"State: state Sketch(KLL k=200)"| m + m -->|"State: state Sketch(KLL k=200)"| out + p0 -.- c0 + p1 -.- c1 + m -.- cm + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +``` + +Both panes have the same `definition` (`SummaryAgg` over `Scan m`), and their selections are adjacent, so they merge into one range: ```text -pane 0 = SummaryAgg(KLL k=200, column(SampleValue), by[]) over TimeRange(1m, TimeShift(0, Scan m)) - coverage() = { definition: SummaryAgg(...) over Scan m, selection: [{ time: Relative (−1m, 0] }] } -pane 1 = SummaryAgg(KLL k=200, column(SampleValue), by[]) over TimeRange(1m, TimeShift(1m, Scan m)) - coverage() = { definition: same, selection: [{ time: Relative (−2m, −1m] }] } -SummaryMerge(children = [pane 0, pane 1], group_by = by[]) - ──State(state Sketch(KLL{k=200}))──▶ - coverage() = { definition: same, selection: [{ time: Relative (−2m, 0] }] } +time −2m −1m 0 +pane 1 (─────────────] +pane 0 (─────────────] +merge (───────────────────────────] ``` -Scenario B, populations: `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1) merge into `selection: [{ t.region: In{'us', 'eu'} }]`. +**Scenario B, populations.** `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1) merge: + +```mermaid +flowchart TB + scan["Scan t"]:::value + fu["Filter region = 'us'"]:::value + fe["Filter region = 'eu'"]:::value + au["SummaryAgg KLL k=200, by job"]:::state + ae["SummaryAgg KLL k=200, by job"]:::state + m["SummaryMerge
group_by: by job"]:::state + out(["next operator"]) + cu["selection: region ∈ {us}"]:::cov + ce["selection: region ∈ {eu}"]:::cov + cm["selection: region ∈ {us, eu}"]:::cov + scan --> fu --> au + scan --> fe --> ae + au --> m + ae --> m + m -->|"State: job Utf8, state Sketch(KLL k=200)"| out + au -.- cu + ae -.- ce + m -.- cm + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +``` -Scenario C, rollup: one `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Every job's state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. +**Scenario C, rollup (planned).** One `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Each job's output state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. + +```mermaid +flowchart TB + a["SummaryAgg KLL k=200
by region, job"]:::state + m["SummaryMerge
group_by: by job"]:::state + out(["next operator"]) + a -->|"State: region Utf8, job Utf8, state Sketch(KLL k=200)"| m + m -->|"State: job Utf8, state Sketch(KLL k=200)"| out + classDef value fill:#dbeafe,stroke:#2563eb,color:#111 + classDef state fill:#fef3c7,stroke:#d97706,color:#111 + classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +``` + +```text +input groups output groups +(us, api) ─┐ +(eu, api) ─┴─ merge ─────────────▶ api +(us, web) ─┐ +(eu, web) ─┴─ merge ─────────────▶ web +``` - Output schema: the children's schema with the group key fields reduced to `group_by`. - Checks: @@ -575,6 +678,15 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | | `Extension` | `child, name` | deployment-named state operator | +`SummarySubtract`, for a family that allows it (e.g. an exact `Sum`): + +```text +time −10m −5m 0 +left (───────────────────────────────────────] +right (───────────────────] +result (───────────────────] +``` + ## 6. Key code interfaces `OperatorNode`, `OperatorResultKind` and coverage are in §4. Bodies and serde/derive attributes are elided below. From 7b52ded457ba3347cd17bae2dd2f4ccdd522ea43 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:19:54 +0000 Subject: [PATCH 26/59] =?UTF-8?q?docs:=20reword=20how=20coverage=20is=20co?= =?UTF-8?q?mputed=20in=20=C2=A74.4?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 261563156..50f3f7342 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -244,7 +244,7 @@ Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` ### 4.4 Deriving the selection -Nobody declares coverage: the planner computes it from the sub-DAG. It starts at the `SummaryAgg`, walks down, and looks at each filter condition on the way (split at `AND`). A condition moves into `selection` only when both rules hold: +The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. It starts at the `SummaryAgg`, walks down, and looks at each filter condition on the way (split at `AND`). A condition moves into `selection` only when both rules hold: - **Rule 1: it can move up to the `SummaryAgg` without changing its meaning.** - **Rule 2: it is a simple condition on one column.** From 1fffe64b3bf08ae0414904b08b5d72ae651672a8 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:21:29 +0000 Subject: [PATCH 27/59] =?UTF-8?q?docs:=20explain=20the=20selection=20deriv?= =?UTF-8?q?ation=20in=20=C2=A74.4=20as=20goal=20and=20steps?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 21 ++++++++++++------- 1 file changed, 14 insertions(+), 7 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 50f3f7342..84e98d546 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -244,12 +244,19 @@ Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` ### 4.4 Deriving the selection -The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. It starts at the `SummaryAgg`, walks down, and looks at each filter condition on the way (split at `AND`). A condition moves into `selection` only when both rules hold: +The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. -- **Rule 1: it can move up to the `SummaryAgg` without changing its meaning.** -- **Rule 2: it is a simple condition on one column.** +**Goal.** Split the filter conditions under a `SummaryAgg` into two groups: conditions that only choose *which rows* go into the state (they go into `selection`), and everything else (it stays in `definition`). -Otherwise it stays in `definition`, like a residual in the paper. +**Steps.** + +1. **Collect the conditions.** Go down the sub-DAG from the `SummaryAgg` and collect every filter condition: from `Filter` nodes, from `Scan.predicates`, and from the `SummaryAgg`'s own `filter`. A condition `A AND B` counts as two conditions, `A` and `B`. +2. **Ask two questions about each condition:** + - **Rule 1: would it pick the same rows if it sat directly under the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. + - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. +3. **Sort it.** + - Both answers yes: take the condition out of the sub-DAG and put it into `selection`. + - Otherwise: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. **Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: @@ -262,7 +269,7 @@ SummaryAgg KLL(value) by[job] filter: value < 100 └─ Scan m ``` -| Condition | Found at | Rule 1: moves up? | Rule 2: simple? | Result | +| Condition | Found at | Rule 1: same rows at the `SummaryAgg`? | Rule 2: simple? | Result | |---|---|---|---|---| | `value < 100` | `SummaryAgg.filter` | yes, it is already there | yes, an interval | `selection`: `value ∈ (−∞, 100)` | | `region = 'us'` | `Filter` | yes: `Project` passes `region` through (renamed `r`) | yes, a value set | `selection`: `m.region ∈ {us}` | @@ -276,9 +283,9 @@ definition: KLL(value) by[job] over Project[job, region AS r, value] over Filter selection: value ∈ (−∞, 100), m.region ∈ {us}, time (−3m, −2m] ``` -**Rule 1 in detail.** This is the reverse of filter pushdown (DataFusion's `PushDownFilter`). A condition can move up past: +**Rule 1 in detail.** The answer depends on the operators between the condition and the `SummaryAgg`. The condition must be able to pass each of them (this is the reverse of filter pushdown, DataFusion's `PushDownFilter`): -| Operator | Moves up? | Example | +| Operator in between | The condition can pass it | Example | |---|---|---| | `Filter`, `Scan.predicates`, `SummaryAgg.filter` | always | above | | range `TimeRange`, `TimeShift` without `@` | always | above | From 8fb77076209431ca802279474d7a3608ed985b91 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:22:11 +0000 Subject: [PATCH 28/59] =?UTF-8?q?docs:=20move=20the=20Goldstein=20&=20Lars?= =?UTF-8?q?on=20mapping=20into=20a=20footnote=20of=20=C2=A74.2?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 48 +++++++------------ 1 file changed, 16 insertions(+), 32 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 84e98d546..483a8c128 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -48,7 +48,7 @@ This section fixes the words used below. They follow relational databases and Ap |---|---|---| | **Relation / table** | A set (bag) of rows with the same columns. A base table is stored; a derived relation is the output of a query operator. | Every edge in the DAG carries a relation. A `Scan` reads a base table (SQL table or PromQL metric); every other operator outputs a derived relation. | | **Row / tuple** | One element of a relation: one value per column. | One output row of a node. For PromQL, one sample of one series at one time. | -| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.4). | +| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.3). | | **Schema** | The ordered list of columns of a relation: name, data type, nullability (Arrow `Schema` of `Field { name, data_type, nullable }`; DataFusion `DFSchema` adds the table qualifier). The schema is *metadata*: it describes rows, it contains none. | `Schema` of `Field { name, dtype, nullable, table }` in `crates/types/src/pre_asap/schema.rs`. Unlike Arrow, `dtype` can be a summary state type (§3). | | **Data type** | The type of a column's values (`Int64`, `Utf8`, `Timestamp`, …). | `DataType`, wrapped as `FieldDataType::Plain`. | | **Aggregate state** | The intermediate value of an aggregate function before its final result, e.g. `(sum, count)` for `AVG` (DataFusion `Accumulator::state`, partial/final aggregation). It is never exposed as a column type to users. | Summary state *is* a column type here (`FieldDataType::Sketch`, `ExactAggregate`, …), so state can flow along edges and be merged, stored and read by later operators. | @@ -113,11 +113,10 @@ A node in the physical data will represent the data or summary instance, so a no | Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | | What is it based on? | Goldstein & Larson view matching (SIGMOD 2001), explained with an example | §4.1 | | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | -| What do we take from the paper, and what do we add? | computation, aggregate and `GROUP BY` → `definition`; ranges → `selection`; plus unions and summary families | §4.3 | -| Who sets it? | Nobody: it is derived from the sub-DAG | §4.4 | -| What uses it? | merge, rollup, slice, reuse, subtract | §4.5 | -| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.6 | -| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.7 | +| How is it computed? | by the planner, from the sub-DAG the node covers | §4.3 | +| What uses it? | merge, rollup, slice, reuse, subtract | §4.4 | +| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.5 | +| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.6 | **Why coverage is not part of the schema.** @@ -193,7 +192,7 @@ GROUP BY job; ### 4.2 Summary Coverage = Summary definition + selection -A summary state is a stored aggregation, like `V` above, whose aggregate is a sketch. So we describe it the way the paper describes a view, in two parts: +A summary state is a stored aggregation, like `V` above, whose aggregate is a sketch. So we describe it the way the paper describes a view, in two parts[^gl]: | Part | Question it answers | What it is | |---|---|---| @@ -221,28 +220,13 @@ selection: region ∈ {us}, latency ∈ (−∞, 100) | `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | | `Filter(value * 2 > 10, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `value * 2 > 10` | nothing | -In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column (§4.4). +In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column (§4.3). Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` is the sub-DAG without its row filters, `σ` is the selection, and `G` the grouping. -### 4.3 What we take from Goldstein & Larson, and what we add +[^gl]: **What we take from Goldstein & Larson, and what we add.** The view's tables, joins and residuals become the sub-DAG `C` below the `SummaryAgg`; the aggregate and its argument become the summary family and its input; `GROUP BY` becomes the `SummaryAgg` grouping `G`. These three are in `definition`. The paper's ranges become `selection` (§4.3). A compensating filter on the view's output becomes slicing, allowed only on a column of `G`, and regrouping to a smaller `GROUP BY` becomes rollup (§4.4). We add three things: **unions of states** (the paper uses one view at a time; `SummaryMerge` combines several, so we also check that their selections do not overlap), **summary families** (the paper only re-adds `SUM` and `COUNT`; each family says how its inputs may overlap, §4.4), and **more kinds of conditions** (value sets, hash partitions, and time relative to the evaluation time, §4.3). -| Goldstein & Larson | Summary coverage | -|---|---| -| tables, joins and residuals of the view | the sub-DAG `C` below the `SummaryAgg`, in `definition` | -| the aggregate and its argument | the summary family and its input, in `definition` | -| `GROUP BY` | the `SummaryAgg` grouping `G`, in `definition` | -| ranges | `selection` (§4.4) | -| compensating filter on the view's output | slicing: allowed only on a column of `G` (§4.5) | -| regrouping to a smaller `GROUP BY` | rollup (§4.5) | - -**What we add:** - -- **Unions of states.** The paper uses one view at a time. `SummaryMerge` combines several states, so we must also check that their selections do not overlap. -- **Summary families.** The paper only re-adds `SUM` and `COUNT`. Each summary family says how its inputs may overlap (§4.5). -- **More kinds of conditions:** value sets (`IN`, `NOT IN`), hash partitions, and time relative to the evaluation time (§4.4). - -### 4.4 Deriving the selection +### 4.3 Deriving the selection The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. @@ -327,7 +311,7 @@ Not simple, so they stay in `definition`: `value * 2 > 10` (expression), `a = b` - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. -### 4.5 Operations +### 4.4 Operations Coverage tells the planner which states can be combined, and what the result covers. The examples below use these states. All are `KLL(value) by[job] over Scan m` unless noted: @@ -377,7 +361,7 @@ Coverage tells the planner which states can be combined, and what the result cov - `SummaryUpdate.weight_domain` is compared too. It is computed from the rest, so it differs only if something is wrong. - States over different tables never merge: a KLL over `m1` and one over `m2` have different definitions. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. -### 4.6 Interface +### 4.5 Interface ```rust pub struct OperatorNode { @@ -429,7 +413,7 @@ impl SummaryCoverage { - `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. - A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. -### 4.7 What coverage does not contain +### 4.6 What coverage does not contain Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to ASAPQuery-backend, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. @@ -474,7 +458,7 @@ flowchart TB - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. - Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **always derived**, never declared (§4.4). Both conditions of the `Filter` move into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. +- Coverage: **always derived**, never declared (§4.3). Both conditions of the `Filter` move into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. - Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. ### 5.2 `SummaryEstimate`: sketch state → value @@ -572,7 +556,7 @@ Current state: - **On `main` (since #560):** `SummaryMerge { children }` is implemented. `validate_inputs()` accepts it when there is at least one child, every child is `State` with exactly one state field, and all children have identical schemas. The output schema is the children's schema. - **#646 (open):** adds the coverage check. `OperatorNode::new` and `validate_structure` also require equal `definition`s and disjoint selections, and `coverage()` returns the merged coverage. -- **Planned:** `group_by`, so one operator does both merge and rollup (§4.5): +- **Planned:** `group_by`, so one operator does both merge and rollup (§4.4): ```rust SummaryMerge { children: Vec, group_by: Reduction } @@ -668,7 +652,7 @@ input groups output groups - Output schema: the children's schema with the group key fields reduced to `group_by`. - Checks: - at least one child, every child is `State` with exactly one state field; - - all children have equal `definition`s (§4.5), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; + - all children have equal `definition`s (§4.4), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; - `group_by` ⊆ the children's `G`, and the family merges; - the children's selections relate as the family requires: disjoint for KLL, so pane 0 with pane 0 is rejected; overlap is allowed for HLL. - Coverage: **derived**: the shared `definition` with `group_by`, and the union of the children's selections. Adjacent intervals join; gaps stay as separate boxes. Nested merges work because a child merge has coverage like any other summary node. @@ -680,7 +664,7 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | Operator | Fields | Intended edge shape | |---|---|---| -| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.5) | +| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.4) | | `SummaryDelete` | `summary_input, key: ColumnId` | State → State with the entries for `key` removed | | `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | | `Extension` | `child, name` | deployment-named state operator | From 063bf6765dee54285b17c50a6e88ccdc0948492e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:23:51 +0000 Subject: [PATCH 29/59] docs: clearer wording for Rule 1 Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 483a8c128..ff79270c8 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -236,7 +236,7 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no 1. **Collect the conditions.** Go down the sub-DAG from the `SummaryAgg` and collect every filter condition: from `Filter` nodes, from `Scan.predicates`, and from the `SummaryAgg`'s own `filter`. A condition `A AND B` counts as two conditions, `A` and `B`. 2. **Ask two questions about each condition:** - - **Rule 1: would it pick the same rows if it sat directly under the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. + - **Rule 1: would it pick the same rows if it were moved to just below the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. 3. **Sort it.** - Both answers yes: take the condition out of the sub-DAG and put it into `selection`. From bcd82164c944a1f543cb9d2c928078f52e9a9a5b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:25:40 +0000 Subject: [PATCH 30/59] =?UTF-8?q?docs:=20draw=20=C2=A75=20examples=20as=20?= =?UTF-8?q?bottom-up=20operator=20trees?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 253 +++++++++--------- 1 file changed, 132 insertions(+), 121 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index ff79270c8..0250b5e92 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -426,13 +426,14 @@ Coverage only says what a state means and which rows it took. The deployment and Given that these information requirements are introduced by summary operators to work correctly semantically, we show the examples of how the defined OperatorNode, schema, and physical data information work with each kind of summary operators. -How to read the diagrams: data flows from top to bottom. Each edge is labelled with the schema it carries, written `Kind: field Type, …`. +How to read the diagrams: data flows from bottom to top, along the `▲` arrows. Each edge is labelled with the schema it carries, written `Kind: field Type, …`. -| Color | Meaning | +| Notation | Meaning | |---|---| -| blue box | operator whose output is a value (`Relation`) | -| yellow box | operator whose output is summary state (`State`) | -| gray dashed box | the node's `coverage()` | +| `[ Op ]` | operator whose output is a value (`Relation`) | +| `[[ Op ]]` | operator whose output is summary state (`State`) | +| `( next operator )` | whatever consumes the result | +| `selection: …` next to a node | the `selection` part of that node's `coverage()` | Schemas are the ones `output_schema()` derives. Planning may rename fields through `OperatorNode::with_schema`, but types, nullability, `time_index`, `unique_keys` and `closed` must match the derivation. All examples use a table source, so values are `Relation`; with a `TimeSeries` source the value side is `InstantVector`. @@ -440,20 +441,26 @@ Schemas are the ones `output_schema()` derives. Planning may rename fields throu Scenario: p99 latency by job, from KLL(k=200), over table `t`, US rows with latency under 10 s only. -```mermaid -flowchart TB - scan["Scan t"]:::value - filter["Filter
region = 'us' AND latency #lt; 10000"]:::value - agg["SummaryAgg
family: Sketch KLL k=200
input: column latency
reduction: by job"]:::state - out(["next operator"]) - cov["coverage()
definition: this SummaryAgg over Scan t (Filter removed)
selection: t.region ∈ {us}, t.latency ∈ (−∞, 10000)"]:::cov - scan -->|"Relation: job Utf8, region Utf8, ts Timestamp, latency Float64"| filter - filter -->|"Relation: same as above"| agg - agg -->|"State: job Utf8, state Sketch(KLL k=200)"| out - agg -.- cov - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ State: job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryAgg ]] KLL k=200, input latency, by job + ▲ + │ Relation: job Utf8, region Utf8, ts Timestamp, latency Float64 + │ + [ Filter ] region = 'us' AND latency < 10000 + ▲ + │ Relation: job Utf8, region Utf8, ts Timestamp, latency Float64 + │ + [ Scan t ] + +coverage() of the SummaryAgg +┌──────────────────────────────────────────────────────────────────┐ +│ definition: this SummaryAgg over Scan t (the Filter removed) │ +│ selection: t.region ∈ {us}, t.latency ∈ (−∞, 10000) │ +└──────────────────────────────────────────────────────────────────┘ ``` - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. @@ -465,16 +472,16 @@ flowchart TB Scenario: read p99 from the state in 5.1. -```mermaid -flowchart TB - agg["SummaryAgg from 5.1"]:::state - est["SummaryEstimate
query: Quantile q = 0.99"]:::value - out(["next operator"]) - agg -->|"State: job Utf8, state Sketch(KLL k=200)"| est - est -->|"Relation: job Utf8, quantile Float64 (planner may rename to p99)"| out - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ Relation: job Utf8, quantile Float64 (planner may rename to p99) + │ + [ SummaryEstimate ] query: Quantile q = 0.99 + ▲ + │ State: job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryAgg ]] from 5.1 ``` - Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`). Keys and metadata pass through. A top-k readout is the exception: it returns the selected rows, one per ranked item, with the partition keys, the item identity columns, and a `value` Float64 score (#579). This is the same row shape as an exact Sort → Limit top-k, so the plans for one query share a root schema. @@ -487,20 +494,26 @@ flowchart TB Scenario: total bytes by host with an exact Sum accumulator. -```mermaid -flowchart TB - scan["Scan t"]:::value - agg["SummaryAgg
family: ExactAggregate Sum
input: column bytes
reduction: by host"]:::state - fin["FinalizeExactAccumulator"]:::value - out(["next operator"]) - cov["coverage()
definition: this SummaryAgg over Scan t
selection: everything"]:::cov - scan -->|"Relation: host Utf8, bytes Float64"| agg - agg -->|"State: host Utf8, state ExactAggregate(Sum)"| fin - fin -->|"Relation: host Utf8, state Float64"| out - agg -.- cov - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ Relation: host Utf8, state Float64 + │ + [ FinalizeExactAccumulator ] + ▲ + │ State: host Utf8, state ExactAggregate(Sum) + │ + [[ SummaryAgg ]] ExactAggregate Sum, input bytes, by host + ▲ + │ Relation: host Utf8, bytes Float64 + │ + [ Scan t ] + +coverage() of the SummaryAgg +┌──────────────────────────────────────────────────┐ +│ definition: this SummaryAgg over Scan t │ +│ selection: everything (no filter) │ +└──────────────────────────────────────────────────┘ ``` - Output schema: each `ExactAggregate` field keeps its name (`state`) and takes the type and nullability the equivalent `NonASAPOp::Aggregate` would give: Sum/Min/Max follow the input column, Count is Int64, and Rate/IRate/Increase are Float64. If the child is not a `SummaryAgg` directly, Count falls back to Int64 and the others to Float64. `unique_keys`, `closed` and `time_index` are preserved (`schema_rebuilding.rs`). @@ -512,16 +525,16 @@ flowchart TB Scenario: keep the full latency population per job, so that p99 and top-10 can be evaluated later. -```mermaid -flowchart TB - scan["Scan t (closed schema)"]:::value - mp["MaintainPopulation
input: Rows of that Scan, value column latency, by job
max_k: 10, quantiles: true"]:::state - out(["next operator"]) - scan -->|"Relation: job Utf8, latency Float64"| mp - mp -->|"State: job Utf8, latency Float64 (all plain)"| out - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ State: job Utf8, latency Float64 (all fields plain) + │ + [[ MaintainPopulation ]] input: rows of that Scan, value latency, by job + ▲ max_k: 10, quantiles: true + │ Relation: job Utf8, latency Float64 + │ + [ Scan t ] closed schema ``` - Output schema: identical to the child's, all plain. Only `result_kind = State` marks it as maintained state. @@ -533,16 +546,16 @@ flowchart TB Scenario: p99 by job from the population in 5.4. -```mermaid -flowchart TB - mp["MaintainPopulation from 5.4"]:::state - ev["EvaluatePopulation
evaluation: Quantile q = 0.99"]:::value - out(["next operator"]) - mp -->|"State: job Utf8, latency Float64"| ev - ev -->|"Relation: job Utf8, quantile_0_99 Float64"| out - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ Relation: job Utf8, quantile_0_99 Float64 + │ + [ EvaluatePopulation ] evaluation: Quantile q = 0.99 + ▲ + │ State: job Utf8, latency Float64 + │ + [[ MaintainPopulation ]] from 5.4 ``` - Output schema: the schema of `Aggregate(by grouping, measure)` over the maintained source. Quantile gives `quantile_` Float64, Sum gives `sum` (value type), Count gives `count` Int64 and Average gives `avg` Float64; `unique_keys = [[0]]`, `closed`. `TopK { k }` instead returns the source schema unchanged (the selected rows). @@ -564,31 +577,31 @@ SummaryMerge { children: Vec, group_by: Reduction } **Scenario A, time panes.** Two one-minute KLL panes of PromQL `quantile_over_time(0.99, m[2m])` merge into the two-minute state. Each pane reads `TimeRange(1m)` over `TimeShift(s)` over the scan, as Stage 2 builds them. Both panes share one `Scan`. -```mermaid -flowchart TB - scan["Scan m"]:::value - s0["TimeShift 0"]:::value - s1["TimeShift 1m"]:::value - r0["TimeRange 1m"]:::value - r1["TimeRange 1m"]:::value - p0["pane 0: SummaryAgg
KLL k=200, by nothing"]:::state - p1["pane 1: SummaryAgg
KLL k=200, by nothing"]:::state - m["SummaryMerge
group_by: nothing"]:::state - out(["next operator"]) - c0["selection: time (−1m, 0]"]:::cov - c1["selection: time (−2m, −1m]"]:::cov - cm["selection: time (−2m, 0]
definition: same as both panes"]:::cov - scan --> s0 --> r0 --> p0 - scan --> s1 --> r1 --> p1 - p0 -->|"State: state Sketch(KLL k=200)"| m - p1 -->|"State: state Sketch(KLL k=200)"| m - m -->|"State: state Sketch(KLL k=200)"| out - p0 -.- c0 - p1 -.- c1 - m -.- cm - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ State: state Sketch(KLL k=200) + │ + [[ SummaryMerge ]] group_by: nothing + ▲ selection: time (−2m, 0] + │ + ┌─────────────────┴─────────────────┐ + │ State: state Sketch(KLL k=200) │ State: state Sketch(KLL k=200) + │ │ + [[ SummaryAgg ]] pane 0 [[ SummaryAgg ]] pane 1 + KLL k=200, by nothing KLL k=200, by nothing + selection: time (−1m, 0] selection: time (−2m, −1m] + ▲ ▲ + │ │ + [ TimeRange 1m ] [ TimeRange 1m ] + ▲ ▲ + │ │ + [ TimeShift 0 ] [ TimeShift 1m ] + ▲ ▲ + │ │ + └─────────────────┬─────────────────┘ + │ + [ Scan m ] ``` Both panes have the same `definition` (`SummaryAgg` over `Scan m`), and their selections are adjacent, so they merge into one range: @@ -602,43 +615,41 @@ merge (────────────────────── **Scenario B, populations.** `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1) merge: -```mermaid -flowchart TB - scan["Scan t"]:::value - fu["Filter region = 'us'"]:::value - fe["Filter region = 'eu'"]:::value - au["SummaryAgg KLL k=200, by job"]:::state - ae["SummaryAgg KLL k=200, by job"]:::state - m["SummaryMerge
group_by: by job"]:::state - out(["next operator"]) - cu["selection: region ∈ {us}"]:::cov - ce["selection: region ∈ {eu}"]:::cov - cm["selection: region ∈ {us, eu}"]:::cov - scan --> fu --> au - scan --> fe --> ae - au --> m - ae --> m - m -->|"State: job Utf8, state Sketch(KLL k=200)"| out - au -.- cu - ae -.- ce - m -.- cm - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ State: job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryMerge ]] group_by: by job + ▲ selection: region ∈ {us, eu} + │ + ┌─────────────────┴─────────────────┐ + │ │ + [[ SummaryAgg ]] [[ SummaryAgg ]] + KLL k=200, by job KLL k=200, by job + selection: region ∈ {us} selection: region ∈ {eu} + ▲ ▲ + │ │ + [ Filter region = 'us' ] [ Filter region = 'eu' ] + ▲ ▲ + │ │ + └─────────────────┬─────────────────┘ + │ + [ Scan t ] ``` **Scenario C, rollup (planned).** One `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Each job's output state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. -```mermaid -flowchart TB - a["SummaryAgg KLL k=200
by region, job"]:::state - m["SummaryMerge
group_by: by job"]:::state - out(["next operator"]) - a -->|"State: region Utf8, job Utf8, state Sketch(KLL k=200)"| m - m -->|"State: job Utf8, state Sketch(KLL k=200)"| out - classDef value fill:#dbeafe,stroke:#2563eb,color:#111 - classDef state fill:#fef3c7,stroke:#d97706,color:#111 - classDef cov fill:#f3f4f6,stroke:#6b7280,stroke-dasharray:4 3,color:#111 +```text + ( next operator ) + ▲ + │ State: job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryMerge ]] group_by: by job + ▲ + │ State: region Utf8, job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryAgg ]] KLL k=200, by region, job ``` ```text From 88069fbd2f7cbb76453076aa5cf22da67c6d4347 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:26:16 +0000 Subject: [PATCH 31/59] =?UTF-8?q?docs:=20draw=20=C2=A74.4=20and=20=C2=A75?= =?UTF-8?q?=20examples=20as=20DAGs=20with=20shared=20nodes?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 116 ++++++++++++------ 1 file changed, 80 insertions(+), 36 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 0250b5e92..d1f464a8d 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -245,12 +245,25 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no **Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: ```text -SummaryAgg KLL(value) by[job] filter: value < 100 -└─ Project [job, region AS r, value] - └─ Filter region = 'us' AND value * 2 > 10 - └─ TimeRange 1m (range) - └─ TimeShift 2m - └─ Scan m + ( next operator ) + ▲ + │ + [[ SummaryAgg ]] KLL(value) by job, filter: value < 100 + ▲ + │ + [ Project ] job, region AS r, value + ▲ + │ + [ Filter ] region = 'us' AND value * 2 > 10 + ▲ + │ + [ TimeRange ] 1m (range) + ▲ + │ + [ TimeShift ] 2m + ▲ + │ + [ Scan m ] ``` | Condition | Found at | Rule 1: same rows at the `SummaryAgg`? | Rule 2: simple? | Result | @@ -470,18 +483,28 @@ coverage() of the SummaryAgg ### 5.2 `SummaryEstimate`: sketch state → value -Scenario: read p99 from the state in 5.1. +Scenario: read p99 and p50 from the state in 5.1. One state feeds both readouts. ```text - ( next operator ) - ▲ - │ Relation: job Utf8, quantile Float64 (planner may rename to p99) - │ - [ SummaryEstimate ] query: Quantile q = 0.99 - ▲ - │ State: job Utf8, state Sketch(KLL k=200) - │ - [[ SummaryAgg ]] from 5.1 + ( next operator ) ( next operator ) + ▲ ▲ + │ Relation: job Utf8, │ Relation: job Utf8, + │ quantile Float64 │ quantile Float64 + │ │ + [ SummaryEstimate ] [ SummaryEstimate ] + Quantile q = 0.99 Quantile q = 0.5 + ▲ ▲ + │ │ + └───────────────────┬───────────────────┘ + │ State: job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryAgg ]] KLL k=200, input latency, by job + ▲ + │ + [ Filter ] region = 'us' AND latency < 10000 + ▲ + │ + [ Scan t ] ``` - Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`). Keys and metadata pass through. A top-k readout is the exception: it returns the selected rows, one per ranked item, with the partition keys, the item identity columns, and a `value` Float64 score (#579). This is the same row shape as an exact Sort → Limit top-k, so the plans for one query share a root schema. @@ -544,18 +567,26 @@ Scenario: keep the full latency population per job, so that p99 and top-10 can b ### 5.5 `EvaluatePopulation`: maintained membership → value -Scenario: p99 by job from the population in 5.4. +Scenario: p99 and the top-10 latencies by job, both from the one population in 5.4. ```text - ( next operator ) - ▲ - │ Relation: job Utf8, quantile_0_99 Float64 - │ - [ EvaluatePopulation ] evaluation: Quantile q = 0.99 - ▲ - │ State: job Utf8, latency Float64 - │ - [[ MaintainPopulation ]] from 5.4 + ( next operator ) ( next operator ) + ▲ ▲ + │ Relation: job Utf8, │ Relation: job Utf8, + │ quantile_0_99 Float64 │ latency Float64 (the top rows) + │ │ + [ EvaluatePopulation ] [ EvaluatePopulation ] + Quantile q = 0.99 TopK k = 10 + ▲ ▲ + │ │ + └───────────────────┬───────────────────┘ + │ State: job Utf8, latency Float64 + │ + [[ MaintainPopulation ]] by job, max_k: 10, quantiles: true + ▲ + │ Relation: job Utf8, latency Float64 + │ + [ Scan t ] closed schema ``` - Output schema: the schema of `Aggregate(by grouping, measure)` over the maintained source. Quantile gives `quantile_` Float64, Sum gives `sum` (value type), Count gives `count` Int64 and Average gives `avg` Float64; `unique_keys = [[0]]`, `closed`. `TopK { k }` instead returns the source schema unchanged (the selected rows). @@ -638,18 +669,31 @@ merge (────────────────────── [ Scan t ] ``` -**Scenario C, rollup (planned).** One `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Each job's output state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. +**Scenario C, rollup (planned).** One `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Each job's output state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. The same `by[region, job]` state also answers p99 per region and job directly, so it feeds two consumers. ```text - ( next operator ) - ▲ - │ State: job Utf8, state Sketch(KLL k=200) - │ - [[ SummaryMerge ]] group_by: by job - ▲ - │ State: region Utf8, job Utf8, state Sketch(KLL k=200) - │ - [[ SummaryAgg ]] KLL k=200, by region, job + ( next operator ) ( next operator ) + ▲ ▲ + │ Relation: job Utf8, │ Relation: region Utf8, job Utf8, + │ quantile Float64 │ quantile Float64 + │ │ + [ SummaryEstimate ] [ SummaryEstimate ] + p99 by job p99 by region, job + ▲ ▲ + │ State: job Utf8, │ + │ state Sketch(KLL k=200) │ + │ │ + [[ SummaryMerge ]] │ + group_by: by job │ + ▲ │ + │ │ + └───────────────────┬───────────────────┘ + │ State: region Utf8, job Utf8, state Sketch(KLL k=200) + │ + [[ SummaryAgg ]] KLL k=200, by region, job + ▲ + │ + [ Scan t ] ``` ```text From bb613140a901782b5ab955121551bb094e6ba60c Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:26:47 +0000 Subject: [PATCH 32/59] =?UTF-8?q?docs:=20rename=20step=203=20in=20=C2=A74.?= =?UTF-8?q?3?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index d1f464a8d..601a5e5bd 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -238,7 +238,7 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no 2. **Ask two questions about each condition:** - **Rule 1: would it pick the same rows if it were moved to just below the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. -3. **Sort it.** +3. **Putting the two rules together.** - Both answers yes: take the condition out of the sub-DAG and put it into `selection`. - Otherwise: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. From cc7baf7110a4712fc3ab2cf901686b9c511140b9 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:27:51 +0000 Subject: [PATCH 33/59] docs: phrase step 3 outcomes as conditions Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 601a5e5bd..cfc8da91b 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -239,8 +239,8 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no - **Rule 1: would it pick the same rows if it were moved to just below the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. 3. **Putting the two rules together.** - - Both answers yes: take the condition out of the sub-DAG and put it into `selection`. - - Otherwise: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. + - If both answers are yes: take the condition out of the sub-DAG and put it into `selection`. + - If either answer is no: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. **Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: From 57fe9a29b9d2599085353d18821a1af21a081fc7 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:34:36 +0000 Subject: [PATCH 34/59] docs: plainer Rule 1 and Rule 2 details with reasons and examples Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 41 +++++++++++-------- 1 file changed, 25 insertions(+), 16 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index cfc8da91b..b4fec2d75 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -280,29 +280,38 @@ definition: KLL(value) by[job] over Project[job, region AS r, value] over Filter selection: value ∈ (−∞, 100), m.region ∈ {us}, time (−3m, −2m] ``` -**Rule 1 in detail.** The answer depends on the operators between the condition and the `SummaryAgg`. The condition must be able to pass each of them (this is the reverse of filter pushdown, DataFusion's `PushDownFilter`): +**Rule 1 in detail.** Imagine moving the condition up, one operator at a time, until it is just below the `SummaryAgg`. Every operator it passes must leave the picked rows unchanged. Whether it can pass depends on what the operator does: -| Operator in between | The condition can pass it | Example | -|---|---|---| -| `Filter`, `Scan.predicates`, `SummaryAgg.filter` | always | above | -| range `TimeRange`, `TimeShift` without `@` | always | above | -| `Project` | only for a column passed through as is (a rename is fine) | `region AS r` ✓; `value * 2 AS v2` ✗ | -| `Aggregate` (later) | only on group columns | below `SUM(value) by[job]`: `job = 'api'` ✓, `value > 5` ✗ | -| window function, or per-series function such as `rate` (later) | only on series labels | below `rate(...)`: `job = 'api'` ✓, `value > 5` ✗ (it would filter raw samples, which changes the rate) | -| anything else | never | | +| Operator it must pass | Can it pass? | Why | Example | +|---|---|---|---| +| another `Filter` (also `Scan.predicates`, `SummaryAgg.filter`) | yes | filters only drop rows, so their order does not matter | `region = 'us'` below `Filter(value < 100)` ✓ | +| `TimeRange` (range) or `TimeShift` (without `@`) | yes | they choose a time window, but do not change any row's values | `region = 'us'` below `TimeRange(1m)` ✓ | +| `Project` | only if its column is passed through unchanged (a rename is fine) | the column must still be there, with the same values, above the `Project` | `Project [job, region AS r]`: `region = 'us'` ✓, it becomes `r = 'us'`. `Project [job, value * 2 AS v2]`: `value > 5` ✗, `value` is gone | +| `Aggregate` (later) | only if it uses group columns | a group column has one value per group, so filtering before or after grouping keeps the same groups | below `SUM(value) by job`: `job = 'api'` ✓; `value > 5` ✗, it changes the sums | +| `rate` or a window function (later) | only if it uses series labels | a label is the same for every sample of a series | below `rate(...)`: `job = 'api'` ✓; `value > 5` ✗, dropping raw samples changes the rate | +| any other operator, e.g. `Join`, `Limit` | no | | | + +A condition *above* `rate` has nothing to pass. `Filter(value > 0, rate(...))` keeps the rate outputs above 0, and those are exactly the rows the `SummaryAgg` reads, so it becomes `value ∈ (0, ∞)` in `selection`. Only the `TimeRange(5m)` under `rate` stays in `definition`. -A condition *above* `rate` is different: it filters the rate outputs, which are exactly the rows the `SummaryAgg` sees. `Filter(job = 'api', rate(...))` gives `job ∈ {api}`, and `Filter(value > 0, rate(...))` gives `value ∈ (0, ∞)` on the rate values. The `TimeRange(5m)` below `rate` stays in `definition` either way. +(This is filter pushdown in reverse. DataFusion's `PushDownFilter` uses the same rules to move filters down.) -**Rule 2 in detail.** A simple condition is one of: +**Rule 2 in detail.** `selection` can hold only two shapes of condition, each on a single column: a set of values, or a range. Hash partitions will be a third shape later. -| Kind | Written as | Example | Becomes | +| Shape | Written as | Example | Stored as | |---|---|---|---| -| value set | `=`, `!=`, `IN`, `NOT IN`, `OR` of `=` on one column | `region IN ('us', 'eu')` | `region ∈ {us, eu}` | -| | | `region != 'test'` | `region ∉ {test}` | -| interval | `<`, `<=`, `>`, `>=` | `value >= 10 AND value < 100` | `value ∈ [10, 100)` | +| allowed values | `=`, `IN`, or `OR` of `=` on the same column | `region IN ('us', 'eu')` | `region ∈ {us, eu}` | +| forbidden values | `!=`, `NOT IN` | `region != 'test'` | `region ∉ {test}` | +| range | `<`, `<=`, `>`, `>=` | `value >= 10 AND value < 100` | `value ∈ [10, 100)` | | hash partition (later) | `hash(columns) mod n = k` | `hash(job) mod 4 = 1` | partition 1 of 4 | -Not simple, so they stay in `definition`: `value * 2 > 10` (expression), `a = b` (two columns), `region = 'us' OR job = 'api'` (two columns), `name LIKE 'web%'` (pattern). +Any other shape stays in `definition`: + +| Condition | Why it is not one of the shapes | +|---|---| +| `value * 2 > 10` | it is on an expression, not a column | +| `a = b` | it compares two columns | +| `region = 'us' OR job = 'api'` | it uses two columns | +| `name LIKE 'web%'` | it is a pattern, not a set of values or a range | **Which column a condition is on.** From 0a2eeb9b8a06d1bf575ac54f6938cbb6b849c3de Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:41:36 +0000 Subject: [PATCH 35/59] =?UTF-8?q?docs:=20=C2=A74.3=20derives=20the=20defin?= =?UTF-8?q?ition=20first,=20then=20the=20selection?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 85 +++++++++++++------ 1 file changed, 59 insertions(+), 26 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index b4fec2d75..d7f684343 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -113,7 +113,7 @@ A node in the physical data will represent the data or summary instance, so a no | Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | | What is it based on? | Goldstein & Larson view matching (SIGMOD 2001), explained with an example | §4.1 | | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | -| How is it computed? | by the planner, from the sub-DAG the node covers | §4.3 | +| How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.3 | | What uses it? | merge, rollup, slice, reuse, subtract | §4.4 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.5 | | What is left out? | the deployment and runtime implementation, e.g. SDS | §4.6 | @@ -226,21 +226,9 @@ Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` [^gl]: **What we take from Goldstein & Larson, and what we add.** The view's tables, joins and residuals become the sub-DAG `C` below the `SummaryAgg`; the aggregate and its argument become the summary family and its input; `GROUP BY` becomes the `SummaryAgg` grouping `G`. These three are in `definition`. The paper's ranges become `selection` (§4.3). A compensating filter on the view's output becomes slicing, allowed only on a column of `G`, and regrouping to a smaller `GROUP BY` becomes rollup (§4.4). We add three things: **unions of states** (the paper uses one view at a time; `SummaryMerge` combines several, so we also check that their selections do not overlap), **summary families** (the paper only re-adds `SUM` and `COUNT`; each family says how its inputs may overlap, §4.4), and **more kinds of conditions** (value sets, hash partitions, and time relative to the evaluation time, §4.3). -### 4.3 Deriving the selection +### 4.3 Deriving the definition and the selection -The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. - -**Goal.** Split the filter conditions under a `SummaryAgg` into two groups: conditions that only choose *which rows* go into the state (they go into `selection`), and everything else (it stays in `definition`). - -**Steps.** - -1. **Collect the conditions.** Go down the sub-DAG from the `SummaryAgg` and collect every filter condition: from `Filter` nodes, from `Scan.predicates`, and from the `SummaryAgg`'s own `filter`. A condition `A AND B` counts as two conditions, `A` and `B`. -2. **Ask two questions about each condition:** - - **Rule 1: would it pick the same rows if it were moved to just below the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. - - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. -3. **Putting the two rules together.** - - If both answers are yes: take the condition out of the sub-DAG and put it into `selection`. - - If either answer is no: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. +The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. One walk down the sub-DAG produces both parts: every filter condition either moves into `selection` or stays in `definition`. §4.3.1 describes what the definition is, and §4.3.2 decides which conditions move. **Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: @@ -266,6 +254,60 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no [ Scan m ] ``` +#### 4.3.1 The definition + +The `definition` is the `SummaryAgg` together with its sub-DAG, with every condition that moves into `selection` (§4.3.2) taken out. Everything else stays exactly as it is: + +| In the sub-DAG | In the `definition` | +|---|---| +| a `Filter` whose conditions all move into `selection` | removed | +| a `Filter` with some conditions that stay | kept, with only the conditions that stay | +| `Scan.predicates` and `SummaryAgg.filter` | trimmed the same way | +| a range `TimeRange` over a `TimeShift` that becomes relative time | removed | +| any other operator | unchanged | + +So the `definition` holds what the state computes: the computation `C` with its remaining conditions, the summary family and its parameters, the input column, and the grouping `G`. + +In the worked example, `value < 100`, `region = 'us'` and the time window move into `selection` (§4.3.2), and `value * 2 > 10` stays: + +```text + ( next operator ) + ▲ + │ + [[ SummaryAgg ]] KLL(value) by job ← filter removed + ▲ + │ + [ Project ] job, region AS r, value ← unchanged + ▲ + │ + [ Filter ] value * 2 > 10 ← region = 'us' removed + ▲ + │ + [ Scan m ] ← TimeRange, TimeShift removed +``` + +**When two definitions are equal.** Merging and reuse (§4.4) require equal definitions. + +- They must have the same structure. Planning details are ignored: `timing`, `guarantee` and `coverage_cache`. So a pane built at ingestion time can merge with one built at query time. +- `SummaryUpdate.weight_domain` is compared too. It is computed from the rest, so it differs only if something is wrong. +- States over different tables never merge: a KLL over `m1` and one over `m2` have different definitions. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. + +#### 4.3.2 The selection + +**Goal.** Decide which filter conditions under the `SummaryAgg` move into `selection`: those that only choose *which rows* go into the state. All other conditions stay in the `definition` (§4.3.1). + +**Steps.** + +1. **Collect the conditions.** Go down the sub-DAG from the `SummaryAgg` and collect every filter condition: from `Filter` nodes, from `Scan.predicates`, and from the `SummaryAgg`'s own `filter`. A condition `A AND B` counts as two conditions, `A` and `B`. +2. **Ask two questions about each condition:** + - **Rule 1: would it pick the same rows if it were moved to just below the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. + - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. +3. **Putting the two rules together.** + - If both answers are yes: take the condition out of the sub-DAG and put it into `selection`. + - If either answer is no: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. + +In the worked example: + | Condition | Found at | Rule 1: same rows at the `SummaryAgg`? | Rule 2: simple? | Result | |---|---|---|---|---| | `value < 100` | `SummaryAgg.filter` | yes, it is already there | yes, an interval | `selection`: `value ∈ (−∞, 100)` | @@ -273,12 +315,7 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no | `value * 2 > 10` | `Filter` | yes | **no**: it is on an expression, not a column | stays in `definition` | | 1 minute, shifted by 2 | `TimeRange` + `TimeShift` | yes | yes, relative time | `selection`: `(−3m, −2m]` | -Result: - -```text -definition: KLL(value) by[job] over Project[job, region AS r, value] over Filter(value * 2 > 10) over Scan m -selection: value ∈ (−∞, 100), m.region ∈ {us}, time (−3m, −2m] -``` +So the `selection` is `value ∈ (−∞, 100)`, `m.region ∈ {us}`, time `(−3m, −2m]`. **Rule 1 in detail.** Imagine moving the condition up, one operator at a time, until it is just below the `SummaryAgg`. Every operator it passes must leave the picked rows unchanged. Whether it can pass depends on what the operator does: @@ -377,11 +414,7 @@ Coverage tells the planner which states can be combined, and what the result cov **Subtract** (`SummarySubtract`, reserved): remove one state from another, for families that allow it (e.g. exact `Sum`/`Count`, Count-Min). A sum over `(−10m, 0]` minus a sum over `(−10m, −5m]` gives `(−5m, 0]`. Allowed when the `definition`s are equal and the right selection is inside the left. -**When two `definition`s are equal.** - -- They must have the same structure. Planning details are ignored: `timing`, `guarantee` and `coverage_cache`. So a pane built at ingestion time can merge with one built at query time. -- `SummaryUpdate.weight_domain` is compared too. It is computed from the rest, so it differs only if something is wrong. -- States over different tables never merge: a KLL over `m1` and one over `m2` have different definitions. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. +Two `definition`s count as equal as described in §4.3.1. ### 4.5 Interface From 5fcb954a6ab9b62a4fac916ef5d89c211fe9b0aa Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:41:54 +0000 Subject: [PATCH 36/59] =?UTF-8?q?docs:=20name=20the=20code=20interface=20i?= =?UTF-8?q?n=20the=20=C2=A74.5=20heading?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index d7f684343..5e622b5bd 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -416,7 +416,7 @@ Coverage tells the planner which states can be combined, and what the result cov Two `definition`s count as equal as described in §4.3.1. -### 4.5 Interface +### 4.5 Code interface: `OperatorNode::coverage()` and `SummaryCoverage` (`crates/types/src/ir/node.rs`, `summary_coverage.rs`) ```rust pub struct OperatorNode { From e42ea972b51e765e182496f1ca308681ee4a981f Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:43:08 +0000 Subject: [PATCH 37/59] =?UTF-8?q?docs:=20move=20the=20coverage=20code=20in?= =?UTF-8?q?terface=20to=20=C2=A76.5?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 112 +++++++++--------- 1 file changed, 57 insertions(+), 55 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 5e622b5bd..1f77b98f9 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -115,8 +115,8 @@ A node in the physical data will represent the data or summary instance, so a no | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | | How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.3 | | What uses it? | merge, rollup, slice, reuse, subtract | §4.4 | -| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §4.5 | -| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.6 | +| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.5 | +| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.5 | **Why coverage is not part of the schema.** @@ -416,59 +416,7 @@ Coverage tells the planner which states can be combined, and what the result cov Two `definition`s count as equal as described in §4.3.1. -### 4.5 Code interface: `OperatorNode::coverage()` and `SummaryCoverage` (`crates/types/src/ir/node.rs`, `summary_coverage.rs`) - -```rust -pub struct OperatorNode { - pub operator: Operator, - pub result_kind: OperatorResultKind, - pub schema: Schema, - pub guarantee: Option, - pub timing: Option, - /// Cache for `coverage()`. Lazily filled, never serialized, ignored by - /// equality, emptied on clone. Not a source of truth: coverage is always - /// re-derivable. - coverage_cache: CoverageCache, -} - -impl OperatorNode { - /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`. - pub fn coverage(&self) -> Option<&SummaryCoverage>; -} - -pub struct SummaryCoverage { - /// The `SummaryAgg` (or rolled-up equivalent) with the selection removed. - pub definition: Rc, - /// Union of boxes over the output rows of the definition's computation. - pub selection: Vec, -} - -pub struct SelectionBox { - pub columns: BTreeMap, // missing column = unrestricted - pub relative_time: Option<(Bound, Bound)>, // ms from evaluation; None = unrestricted -} - -pub struct ColumnIdentity { - pub table: Option, - pub name: String, -} - -pub enum Constraint { - In(Vec), // ScalarValue has no total order (Float64) - NotIn(Vec), - Interval { lower: Bound, upper: Bound }, - // HashPartition { columns, of, index }: added with its first producer. -} - -impl SummaryCoverage { - pub fn derive(node: &OperatorNode) -> Result; -} -``` - -- `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. -- A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. - -### 4.6 What coverage does not contain +### 4.5 What coverage does not contain Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to ASAPQuery-backend, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. @@ -963,3 +911,57 @@ impl ASAPOp { pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>; } ``` + +### 6.5 Summary coverage (`crates/types/src/ir/node.rs`, `crates/types/src/ir/summary_coverage.rs`) + +The code for §4. + +```rust +pub struct OperatorNode { + pub operator: Operator, + pub result_kind: OperatorResultKind, + pub schema: Schema, + pub guarantee: Option, + pub timing: Option, + /// Cache for `coverage()`. Lazily filled, never serialized, ignored by + /// equality, emptied on clone. Not a source of truth: coverage is always + /// re-derivable. + coverage_cache: CoverageCache, +} + +impl OperatorNode { + /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`. + pub fn coverage(&self) -> Option<&SummaryCoverage>; +} + +pub struct SummaryCoverage { + /// The `SummaryAgg` (or rolled-up equivalent) with the selection removed. + pub definition: Rc, + /// Union of boxes over the output rows of the definition's computation. + pub selection: Vec, +} + +pub struct SelectionBox { + pub columns: BTreeMap, // missing column = unrestricted + pub relative_time: Option<(Bound, Bound)>, // ms from evaluation; None = unrestricted +} + +pub struct ColumnIdentity { + pub table: Option, + pub name: String, +} + +pub enum Constraint { + In(Vec), // ScalarValue has no total order (Float64) + NotIn(Vec), + Interval { lower: Bound, upper: Bound }, + // HashPartition { columns, of, index }: added with its first producer. +} + +impl SummaryCoverage { + pub fn derive(node: &OperatorNode) -> Result; +} +``` + +- `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. +- A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. From 47e96770a19e4f5adbf0a254e08a6ec58e315606 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:43:15 +0000 Subject: [PATCH 38/59] =?UTF-8?q?docs:=20=C2=A76=20intro=20points=20to=20?= =?UTF-8?q?=C2=A76.5?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 1f77b98f9..2e9a8198e 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -725,7 +725,7 @@ result (──────────────── ## 6. Key code interfaces -`OperatorNode`, `OperatorResultKind` and coverage are in §4. Bodies and serde/derive attributes are elided below. +`OperatorNode` and coverage are in §6.5. Bodies and serde/derive attributes are elided below. ### 6.1 Schema and field types (`crates/types/src/pre_asap/schema.rs`) From 41cd0b3c38c825b3756bf73773dd53fd4ddaa35a Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 19:46:35 +0000 Subject: [PATCH 39/59] =?UTF-8?q?docs:=20merge=20=C2=A74.2=20and=20=C2=A74?= =?UTF-8?q?.3=20(why=20two=20parts,=20then=20derivation);=20fold=20=C2=A74?= =?UTF-8?q?.4=20into=20=C2=A75.6?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 161 +++++++++--------- 1 file changed, 78 insertions(+), 83 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 2e9a8198e..775fee027 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -48,7 +48,7 @@ This section fixes the words used below. They follow relational databases and Ap |---|---|---| | **Relation / table** | A set (bag) of rows with the same columns. A base table is stored; a derived relation is the output of a query operator. | Every edge in the DAG carries a relation. A `Scan` reads a base table (SQL table or PromQL metric); every other operator outputs a derived relation. | | **Row / tuple** | One element of a relation: one value per column. | One output row of a node. For PromQL, one sample of one series at one time. | -| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.3). | +| **Column** | One position in every row, with a name and a type. Qualified as `table.column` when names can collide (DataFusion `Column { relation, name }`). | `ColumnId` refers to a column of the input schema; `(table, name)` identifies it across nodes (§4.2.2). | | **Schema** | The ordered list of columns of a relation: name, data type, nullability (Arrow `Schema` of `Field { name, data_type, nullable }`; DataFusion `DFSchema` adds the table qualifier). The schema is *metadata*: it describes rows, it contains none. | `Schema` of `Field { name, dtype, nullable, table }` in `crates/types/src/pre_asap/schema.rs`. Unlike Arrow, `dtype` can be a summary state type (§3). | | **Data type** | The type of a column's values (`Int64`, `Utf8`, `Timestamp`, …). | `DataType`, wrapped as `FieldDataType::Plain`. | | **Aggregate state** | The intermediate value of an aggregate function before its final result, e.g. `(sum, count)` for `AVG` (DataFusion `Accumulator::state`, partial/final aggregation). It is never exposed as a column type to users. | Summary state *is* a column type here (`FieldDataType::Sketch`, `ExactAggregate`, …), so state can flow along edges and be merged, stored and read by later operators. | @@ -112,11 +112,11 @@ A node in the physical data will represent the data or summary instance, so a no |---|---|---| | Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | | What is it based on? | Goldstein & Larson view matching (SIGMOD 2001), explained with an example | §4.1 | -| What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2 | -| How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.3 | -| What uses it? | merge, rollup, slice, reuse, subtract | §4.4 | +| What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2.1 | +| How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.2.2 | +| What uses it? | merge, rollup, slice, reuse, subtract | §5.6 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.5 | -| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.5 | +| What is left out? | the deployment and runtime implementation, e.g. SDS | §4.3 | **Why coverage is not part of the schema.** @@ -192,43 +192,39 @@ GROUP BY job; ### 4.2 Summary Coverage = Summary definition + selection -A summary state is a stored aggregation, like `V` above, whose aggregate is a sketch. So we describe it the way the paper describes a view, in two parts[^gl]: +#### 4.2.1 Why coverage has two parts -| Part | Question it answers | What it is | -|---|---|---| -| **`definition`** | *What* is computed? | the `SummaryAgg` node with its row filters taken out: the sub-DAG below it, the summary family and parameters, its input column, and its `GROUP BY` | -| **`selection`** | *Which rows* went in? | the row filters that were taken out, as simple conditions on columns | - -**Example:** +A summary state is a stored aggregation, like `V` in §4.1, whose aggregate is a sketch. Before merging or reusing a state, the planner must answer two different questions about it[^gl]: -```text -state = KLL(latency) by[job] over Filter(region = 'us' AND latency < 100, Scan t) +| Part | Question it answers | +|---|---| +| **`definition`** | *What* is computed? | +| **`selection`** | *Which rows* went in? | -definition: KLL(latency) by[job] over Scan t -selection: region ∈ {us}, latency ∈ (−∞, 100) -``` +**Example.** Three states over table `t`, all with the same schema `(job Utf8, state Sketch(KLL k=200))`: -**Why this is enough.** Two states with the same `definition` come from the same computation. If their selections do not overlap, no row is in both, so merging them counts every row once. This holds whatever the computation contains (joins, unions, `rate`, dedup), so we need no special rule per operator. +| State | Sub-DAG | `definition` | `selection` | +|---|---|---|---| +| `S_us` | `KLL(latency) by[job]` over `Filter(region = 'us', Scan t)` | `KLL(latency) by[job]` over `Scan t` | `region ∈ {us}` | +| `S_eu` | `KLL(latency) by[job]` over `Filter(region = 'eu', Scan t)` | `KLL(latency) by[job]` over `Scan t` | `region ∈ {eu}` | +| `S_size` | `KLL(size) by[job]` over `Filter(region = 'eu', Scan t)` | `KLL(size) by[job]` over `Scan t` | `region ∈ {eu}` | -**More examples** of what goes where: +- `S_us + S_eu` ✓: same computation, different rows. The result is the KLL of latency by job for US and EU, with every row counted once. +- `S_us + S_us` ✗: same computation, same rows. Every US row would be counted twice. +- `S_eu + S_size` ✗: same rows, different computation. One summarizes `latency`, the other `size`; the schema cannot tell them apart. -| Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | -|---|---|---| -| `Filter(region = 'us', Scan t)` | `Scan t` | `region ∈ {us}` | -| `Filter(latency < 100, Scan t)` | `Scan t` | `latency ∈ (−∞, 100)` | -| `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | the last 3 to 2 minutes before evaluation, `(−3m, −2m]` | -| `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | -| `Filter(value * 2 > 10, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `value * 2 > 10` | nothing | +If coverage were one thing, for example the whole sub-DAG compared as a unit, `S_us` and `S_eu` would look different and could never merge. With two parts, the planner checks each question on its own: the definitions must be equal, and the selections must not overlap. -In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column (§4.3). +**Why this is enough.** Two states with the same `definition` come from the same computation. If their selections do not overlap, no row is in both, so merging them counts every row once. This holds whatever the computation contains (joins, unions, `rate`, dedup), so we need no special rule per operator. Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` is the sub-DAG without its row filters, `σ` is the selection, and `G` the grouping. -[^gl]: **What we take from Goldstein & Larson, and what we add.** The view's tables, joins and residuals become the sub-DAG `C` below the `SummaryAgg`; the aggregate and its argument become the summary family and its input; `GROUP BY` becomes the `SummaryAgg` grouping `G`. These three are in `definition`. The paper's ranges become `selection` (§4.3). A compensating filter on the view's output becomes slicing, allowed only on a column of `G`, and regrouping to a smaller `GROUP BY` becomes rollup (§4.4). We add three things: **unions of states** (the paper uses one view at a time; `SummaryMerge` combines several, so we also check that their selections do not overlap), **summary families** (the paper only re-adds `SUM` and `COUNT`; each family says how its inputs may overlap, §4.4), and **more kinds of conditions** (value sets, hash partitions, and time relative to the evaluation time, §4.3). +[^gl]: **What we take from Goldstein & Larson, and what we add.** The view's tables, joins and residuals become the sub-DAG `C` below the `SummaryAgg`; the aggregate and its argument become the summary family and its input; `GROUP BY` becomes the `SummaryAgg` grouping `G`. These three are in `definition`. The paper's ranges become `selection` (§4.2.2). A compensating filter on the view's output becomes slicing, allowed only on a column of `G`, and regrouping to a smaller `GROUP BY` becomes rollup (§5.6). We add three things: **unions of states** (the paper uses one view at a time; `SummaryMerge` combines several, so we also check that their selections do not overlap), **summary families** (the paper only re-adds `SUM` and `COUNT`; each family says how its inputs may overlap, §5.6), and **more kinds of conditions** (value sets, hash partitions, and time relative to the evaluation time, §4.2.2). -### 4.3 Deriving the definition and the selection -The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. One walk down the sub-DAG produces both parts: every filter condition either moves into `selection` or stays in `definition`. §4.3.1 describes what the definition is, and §4.3.2 decides which conditions move. +#### 4.2.2 Deriving the definition and the selection + +The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. One walk down the sub-DAG produces both parts: every filter condition either moves into `selection` or stays in `definition`. The definition is described first, then how the planner decides which conditions move into the selection. **Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: @@ -254,9 +250,10 @@ The planner computes the coverage of a node from the sub-DAG the node covers, no [ Scan m ] ``` -#### 4.3.1 The definition +##### The definition + -The `definition` is the `SummaryAgg` together with its sub-DAG, with every condition that moves into `selection` (§4.3.2) taken out. Everything else stays exactly as it is: +The `definition` is the `SummaryAgg` together with its sub-DAG, with every condition that moves into `selection` (below) taken out. Everything else stays exactly as it is: | In the sub-DAG | In the `definition` | |---|---| @@ -268,7 +265,7 @@ The `definition` is the `SummaryAgg` together with its sub-DAG, with every condi So the `definition` holds what the state computes: the computation `C` with its remaining conditions, the summary family and its parameters, the input column, and the grouping `G`. -In the worked example, `value < 100`, `region = 'us'` and the time window move into `selection` (§4.3.2), and `value * 2 > 10` stays: +In the worked example, `value < 100`, `region = 'us'` and the time window move into `selection` (below), and `value * 2 > 10` stays: ```text ( next operator ) @@ -286,15 +283,16 @@ In the worked example, `value < 100`, `region = 'us'` and the time window move i [ Scan m ] ← TimeRange, TimeShift removed ``` -**When two definitions are equal.** Merging and reuse (§4.4) require equal definitions. +**When two definitions are equal.** Merging (§5.6) requires equal definitions. - They must have the same structure. Planning details are ignored: `timing`, `guarantee` and `coverage_cache`. So a pane built at ingestion time can merge with one built at query time. - `SummaryUpdate.weight_domain` is compared too. It is computed from the rest, so it differs only if something is wrong. - States over different tables never merge: a KLL over `m1` and one over `m2` have different definitions. To combine tables, put a `UNION ALL` with a column that marks the source table below one `SummaryAgg`; that column can then be used in `selection` or in the grouping. -#### 4.3.2 The selection +##### The selection -**Goal.** Decide which filter conditions under the `SummaryAgg` move into `selection`: those that only choose *which rows* go into the state. All other conditions stay in the `definition` (§4.3.1). + +**Goal.** Decide which filter conditions under the `SummaryAgg` move into `selection`: those that only choose *which rows* go into the state. All other conditions stay in the `definition` (above). **Steps.** @@ -370,53 +368,19 @@ Any other shape stays in `definition`: - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. -### 4.4 Operations - -Coverage tells the planner which states can be combined, and what the result covers. The examples below use these states. All are `KLL(value) by[job] over Scan m` unless noted: - -| State | Selection | -|---|---| -| `A` | time `(−1m, 0]` | -| `B` | time `(−2m, −1m]` | -| `C` | time `(−90s, −30s]` | -| `D` | time `(−1m, 0]`, but the `definition` has the residual `value * 2 > 10` | -| `E` | time `(−1m, 0]`, `region ∈ {us}` | -| `F` | time `(−1m, 0]`, `region ∈ {eu}` | - -**Merge** (`SummaryMerge`): combine states into one. Allowed when all `definition`s are equal and the selections relate as the family requires. - -| Merge | Allowed? | Why | Result's selection | -|---|---|---|---| -| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (adjacent ranges join) | -| `E + F` | ✓ | same definition, `us` and `eu` do not overlap | `(−1m, 0]`, `region ∈ {us, eu}` | -| `A + C` | ✗ | `(−60s, −30s]` is in both: those rows would be counted twice | | -| `A + A` | ✗ | every row is in both | | -| `A + D` | ✗ | different definitions: `D` only has rows with `value * 2 > 10` | | -| `A + B'` where `B'` is KLL with `k = 400` | ✗ | different definitions (parameters) | | -| `(A + B) + B''` where `B''` covers `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | - -**How inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): +**More examples** of what goes where: -| Rule | Families | Example | +| Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | |---|---|---| -| **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | -| **may overlap** | HLL, exact `Min`/`Max`, distinct sets | HLL over `A`'s and `C`'s selections ✓: a value seen twice is still one distinct value; the result covers `(−90s, 0]` | -| **right inside left** | subtraction | see subtract below | - -**Rollup** (`SummaryMerge` with `group_by`, later): make the grouping coarser. A `by[region, job]` state rolls up to `by[job]`: the state for `job = api` is the merge of `(us, api)`, `(eu, api)`, …. No overlap check is needed, because a row has one `region` and so is in only one group. The selection is unchanged. - -**Slice**: read only some groups. From a `by[region, job]` state, a query for `region = 'us'` by job reads the groups with `region = us` ✓. A query for `value < 50` ✗: `value` is not a grouping column, and a sketch cannot be filtered after it is built. - -**Reuse** for a query: a stored state answers a query when the `definition`s are equal and the query's rows are all in the state, with any difference covered by a slice. A stored `by[region, job]` state over `(−5m, 0]`: - -- p99 by job over the last 5 minutes for `region = 'us'`: ✓ (slice on `region`). -- p99 by job over the last 1 minute: ✗. The state also holds minutes 2–5, and time is not a grouping column, so they cannot be taken out. - -**Subtract** (`SummarySubtract`, reserved): remove one state from another, for families that allow it (e.g. exact `Sum`/`Count`, Count-Min). A sum over `(−10m, 0]` minus a sum over `(−10m, −5m]` gives `(−5m, 0]`. Allowed when the `definition`s are equal and the right selection is inside the left. +| `Filter(region = 'us', Scan t)` | `Scan t` | `region ∈ {us}` | +| `Filter(latency < 100, Scan t)` | `Scan t` | `latency ∈ (−∞, 100)` | +| `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | the last 3 to 2 minutes before evaluation, `(−3m, −2m]` | +| `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | +| `Filter(value * 2 > 10, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `value * 2 > 10` | nothing | -Two `definition`s count as equal as described in §4.3.1. +In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column. -### 4.5 What coverage does not contain +### 4.3 What coverage does not contain Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to ASAPQuery-backend, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. @@ -468,7 +432,7 @@ coverage() of the SummaryAgg - Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. - Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **always derived**, never declared (§4.3). Both conditions of the `Filter` move into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. +- Coverage: **always derived**, never declared (§4.2.2). Both conditions of the `Filter` move into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. - Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. ### 5.2 `SummaryEstimate`: sketch state → value @@ -590,7 +554,7 @@ Current state: - **On `main` (since #560):** `SummaryMerge { children }` is implemented. `validate_inputs()` accepts it when there is at least one child, every child is `State` with exactly one state field, and all children have identical schemas. The output schema is the children's schema. - **#646 (open):** adds the coverage check. `OperatorNode::new` and `validate_structure` also require equal `definition`s and disjoint selections, and `coverage()` returns the merged coverage. -- **Planned:** `group_by`, so one operator does both merge and rollup (§4.4): +- **Planned:** `group_by`, so one operator does both merge and rollup: ```rust SummaryMerge { children: Vec, group_by: Reduction } @@ -694,10 +658,41 @@ input groups output groups (eu, web) ─┴─ merge ─────────────▶ web ``` +**Which merges are allowed.** All states below are `KLL(value) by[job] over Scan m` unless noted: + +| State | Selection | +|---|---| +| `A` | time `(−1m, 0]` | +| `B` | time `(−2m, −1m]` | +| `C` | time `(−90s, −30s]` | +| `D` | time `(−1m, 0]`, but the `definition` keeps `value * 2 > 10` | + +| Merge | Allowed? | Why | Result's selection | +|---|---|---|---| +| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (adjacent ranges join) | +| `A + C` | ✗ | `(−60s, −30s]` is in both, so those rows would be counted twice | | +| `A + A` | ✗ | every row is in both | | +| `A + D` | ✗ | different definitions: `D` only has rows with `value * 2 > 10` | | +| `A` + a KLL with `k = 400` | ✗ | different definitions (parameters) | | +| `(A + B)` + a state over `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | + +**How inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): + +| Rule | Families | Example | +|---|---|---| +| **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | +| **may overlap** | HLL, exact `Min`/`Max`, distinct sets | HLL over `A`'s and `C`'s rows ✓: a value seen twice is still one distinct value; the result covers `(−90s, 0]` | +| **right inside left** | subtraction | §5.7 | + +**Coverage outside `SummaryMerge`.** The planner also uses coverage to read or reuse a state: + +- **Slice:** from a `by[region, job]` state, a query for `region = 'us'` by job reads only the `us` groups ✓. A query for `value < 50` ✗: `value` is not a grouping column, and a sketch cannot be filtered after it is built. +- **Reuse:** a stored state answers a query when the definitions are equal and the query's rows are all in the state, with any difference covered by a slice. A stored `by[region, job]` state over `(−5m, 0]` answers p99 by job over the last 5 minutes for `region = 'us'` ✓, but not over the last 1 minute ✗: time is not a grouping column, so minutes 2–5 cannot be taken out. + - Output schema: the children's schema with the group key fields reduced to `group_by`. - Checks: - at least one child, every child is `State` with exactly one state field; - - all children have equal `definition`s (§4.4), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; + - all children have equal `definition`s (§4.2.2), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; - `group_by` ⊆ the children's `G`, and the family merges; - the children's selections relate as the family requires: disjoint for KLL, so pane 0 with pane 0 is rejected; overlap is allowed for HLL. - Coverage: **derived**: the shared `definition` with `group_by`, and the union of the children's selections. Adjacent intervals join; gaps stay as separate boxes. Nested merges work because a child merge has coverage like any other summary node. @@ -709,7 +704,7 @@ These variants exist so that plans can name them, but `output_schema()`/`validat | Operator | Fields | Intended edge shape | |---|---|---| -| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.4) | +| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.2) | | `SummaryDelete` | `summary_input, key: ColumnId` | State → State with the entries for `key` removed | | `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | | `Extension` | `child, name` | deployment-named state operator | From 44c39a434281a513f8d69f402cc38601fa890079 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:34:14 +0000 Subject: [PATCH 40/59] docs: refer to the downstream deployment runtime, not a specific backend Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 775fee027..570ea54eb 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -382,11 +382,11 @@ In the last two rows `TimeRange(5m)` stays in `definition`: it is the input wind ### 4.3 What coverage does not contain -Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to ASAPQuery-backend, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. +Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to the downstream deployment runtime, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. **SDS mapping.** The SDS split matches coverage: -- `SummaryDefinition` stores the serialized `definition`. Planner provides its serde; Backend owns the format version, definition id and hash. +- `SummaryDefinition` stores the serialized `definition`. Planner provides its serde; the downstream deployment runtime owns the format version, definition id and hash. - A `StoredSummary`'s coordinates are the `selection` bound to one evaluation, plus the group value. ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators From 03eca9990f3d6beed8d910e85ee649fde1171bed Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:34:22 +0000 Subject: [PATCH 41/59] =?UTF-8?q?docs:=20same=20wording=20in=20the=20?= =?UTF-8?q?=C2=A72=20layout=20row?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 570ea54eb..be8d2e28b 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -53,7 +53,7 @@ This section fixes the words used below. They follow relational databases and Ap | **Data type** | The type of a column's values (`Int64`, `Utf8`, `Timestamp`, …). | `DataType`, wrapped as `FieldDataType::Plain`. | | **Aggregate state** | The intermediate value of an aggregate function before its final result, e.g. `(sum, count)` for `AVG` (DataFusion `Accumulator::state`, partial/final aggregation). It is never exposed as a column type to users. | Summary state *is* a column type here (`FieldDataType::Sketch`, `ExactAggregate`, …), so state can flow along edges and be merged, stored and read by later operators. | | **View / materialized view** | A view is a named query (its *definition*). A materialized view also stores the query's result rows; a query can then be answered from it when its definition matches (view matching, §4.1). | A built summary state is a materialized aggregation view whose aggregate is a summary family. Its definition and which rows it took are its coverage (§4). | -| **Physical data layout** | How rows are stored: row-oriented or columnar (Arrow `RecordBatch`: one array per column), split into partitions (hash or range) and batches. | Decided in physical planning ([planning stages §2](planner-layering.md#2-physical-asap-aware-optimization)) and by the executing backend. The logical schema does not depend on it. | +| **Physical data layout** | How rows are stored: row-oriented or columnar (Arrow `RecordBatch`: one array per column), split into partitions (hash or range) and batches. | Decided in physical planning ([planning stages §2](planner-layering.md#2-physical-asap-aware-optimization)) and by the downstream deployment runtime. The logical schema does not depend on it. | Two consequences for the design: From 66e1bce1add5bdb6821a7a8e385d42e93891e5fe Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:38:45 +0000 Subject: [PATCH 42/59] docs: leave the SDS definition as a TODO for a separate doc Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 5 +---- 1 file changed, 1 insertion(+), 4 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index be8d2e28b..835b5e2f2 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -384,10 +384,7 @@ In the last two rows `TimeRange(5m)` stays in `definition`: it is the input wind Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to the downstream deployment runtime, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. -**SDS mapping.** The SDS split matches coverage: - -- `SummaryDefinition` stores the serialized `definition`. Planner provides its serde; the downstream deployment runtime owns the format version, definition id and hash. -- A `StoredSummary`'s coordinates are the `selection` bound to one evaluation, plus the group value. +TODO: the SDS definition, including how it stores a summary's `definition` and `selection`, will be specified in a separate doc. ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators From b94402c3310c3e04675cd271521ec44d58167e21 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:40:48 +0000 Subject: [PATCH 43/59] =?UTF-8?q?docs:=20make=20=C2=A75=20easier=20to=20re?= =?UTF-8?q?ad=20with=20one=20question=20table=20per=20operator?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 138 +++++++++++------- 1 file changed, 82 insertions(+), 56 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 835b5e2f2..20e6f4818 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -388,9 +388,16 @@ TODO: the SDS definition, including how it stores a summary's `definition` and ` ## 5. Examples on how OperatorNode, schema, and physical data information are being used with Summary operators -Given that these information requirements are introduced by summary operators to work correctly semantically, we show the examples of how the defined OperatorNode, schema, and physical data information work with each kind of summary operators. +This section walks through each summary operator with one small example. For each operator it answers four questions: -How to read the diagrams: data flows from bottom to top, along the `▲` arrows. Each edge is labelled with the schema it carries, written `Kind: field Type, …`. +| Question | What it tells you | +|---|---| +| **What comes out?** | the output schema the planner derives (`output_schema()`) | +| **When is it rejected?** | the checks the planner runs when it builds the node | +| **What is its coverage?** | `coverage()` of the output (§4.2) | +| **State or value?** | whether the output is summary state or a readable value | + +**How to read the diagrams.** Data flows from bottom to top, along the `▲` arrows. Each edge is labelled with the schema it carries, written `Kind: field Type, …`. | Notation | Meaning | |---|---| @@ -399,11 +406,13 @@ How to read the diagrams: data flows from bottom to top, along the `▲` arrows. | `( next operator )` | whatever consumes the result | | `selection: …` next to a node | the `selection` part of that node's `coverage()` | -Schemas are the ones `output_schema()` derives. Planning may rename fields through `OperatorNode::with_schema`, but types, nullability, `time_index`, `unique_keys` and `closed` must match the derivation. All examples use a table source, so values are `Relation`; with a `TimeSeries` source the value side is `InstantVector`. +All examples read a table, so values are `Relation`. For PromQL series they would be `InstantVector`. ### 5.1 `SummaryAgg`: values → state -Scenario: p99 latency by job, from KLL(k=200), over table `t`, US rows with latency under 10 s only. +**What it does.** Turns rows into summary state: one state per group. + +**Example.** p99 latency by job, from a KLL sketch with `k = 200`, over table `t`, using only US rows with latency under 10 s. ```text ( next operator ) @@ -427,14 +436,18 @@ coverage() of the SummaryAgg └──────────────────────────────────────────────────────────────────┘ ``` -- Output schema: the `by` keys followed by one non-nullable field `state` typed `family`; `unique_keys = [[0]]`, `closed = true`, no `time_index`. With `Reduction::PerEntity` the input columns are kept and the sample-value column is replaced by `state`. -- Checks: `family` is not `Plain`; the child is not `State`; the `weight`/`item` columns resolve against the child schema; `filter`, if present, types as `Bool`. -- Coverage: **always derived**, never declared (§4.2.2). Both conditions of the `Filter` move into `selection`, so `definition` is this node over the bare `Scan`. A KLL over `latency` for `region = 'eu'` has the same `definition` and a disjoint selection, so the two can merge. A conjunct that cannot lift (say `latency * 2 > 10`) stays in `definition` as a residual; the node still has coverage. -- Boundary: this is where values become state. The sketch family, algorithm and parameters are committed in the field type, and `guarantee` stays `None` because state is not a caller-visible value. +| Question | Answer | +|---|---| +| What comes out? | the group columns (`job`), then one field `state` whose type is the summary type, here `Sketch(KLL k=200)`. Each `job` appears once | +| When is it rejected? | the summary type is a plain value type; the input is already state; the input column (`latency`) is not in the child's schema; `filter` is not a boolean | +| What is its coverage? | always present. Both `Filter` conditions are simple, so they move into `selection`, and the `definition` is the `SummaryAgg` over the bare `Scan t`. A condition like `latency * 2 > 10` would stay in the `definition` (§4.2.2) | +| State or value? | state. This is where values become state, so the result has no error bound yet | ### 5.2 `SummaryEstimate`: sketch state → value -Scenario: read p99 and p50 from the state in 5.1. One state feeds both readouts. +**What it does.** Reads a number out of a sketch, for example a quantile or a count. + +**Example.** Read p99 and p50 from the state in 5.1. One state feeds both readouts. ```text ( next operator ) ( next operator ) @@ -458,15 +471,18 @@ Scenario: read p99 and p50 from the state in 5.1. One state feeds both readouts. [ Scan t ] ``` -- Output schema: the input schema with the one non-plain field replaced by a non-nullable plain field. Its name and type come from the statistic: `quantile`/`frequency_l2`/`frequency_entropy` Float64, `cardinality`/`count` Int64 (Float64 if the producer is a `PerEntity` `SummaryAgg`). Keys and metadata pass through. A top-k readout is the exception: it returns the selected rows, one per ranked item, with the partition keys, the item identity columns, and a `value` Float64 score (#579). This is the same row shape as an exact Sort → Limit top-k, so the plans for one query share a root schema. -- Result kind: the value kind of the source the state was built from (`Relation` here). -- Checks: input is `State` with exactly one non-plain field, that field is `Sketch`, and its category accepts the statistic (§3). For example, `Cardinality` on KLL is rejected. -- Coverage: **none**. The output is a value; `coverage()` returns `None`. -- Boundary: state is consumed and a value is produced; `guarantee` on this node carries the readout's error bound. +| Question | Answer | +|---|---| +| What comes out? | the same columns, with `state` replaced by the answer: `quantile` Float64 here. Counts and cardinalities are Int64. A top-k readout instead returns the top rows themselves (the same shape as an exact `Sort` + `Limit`) | +| When is it rejected? | the input is not a sketch, or the sketch cannot answer the question. For example, asking a KLL for a cardinality | +| What is its coverage? | none: the output is a value | +| State or value? | value. The node carries the readout's error bound | ### 5.3 `FinalizeExactAccumulator`: exact state → value -Scenario: total bytes by host with an exact Sum accumulator. +**What it does.** Turns an exact accumulator (sum, count, min, max, rate, …) into its final value. + +**Example.** Total bytes by host, with an exact `Sum` accumulator. ```text ( next operator ) @@ -490,14 +506,18 @@ coverage() of the SummaryAgg └──────────────────────────────────────────────────┘ ``` -- Output schema: each `ExactAggregate` field keeps its name (`state`) and takes the type and nullability the equivalent `NonASAPOp::Aggregate` would give: Sum/Min/Max follow the input column, Count is Int64, and Rate/IRate/Increase are Float64. If the child is not a `SummaryAgg` directly, Count falls back to Int64 and the others to Float64. `unique_keys`, `closed` and `time_index` are preserved (`schema_rebuilding.rs`). -- Checks: the input is `State` and contains an `ExactAggregate` field; a sketch is rejected (`structure_contract.rs`). -- Coverage: **none** on the output. -- Boundary: this is the explicit maintenance-to-read boundary for exact state. Exact state is never read through `SummaryEstimate`. +| Question | Answer | +|---|---| +| What comes out? | the same columns, with `state` turned into the value type an ordinary `Aggregate` would give: a sum of Float64 is Float64, a count is Int64, a rate is Float64 | +| When is it rejected? | the input has no exact accumulator, for example a sketch (sketches are read with `SummaryEstimate`) | +| What is its coverage? | none on the output. The `SummaryAgg` below has coverage as in 5.1, whose selection has no conditions (all rows), because there is no filter | +| State or value? | value. Exact state is only ever read through this operator | ### 5.4 `MaintainPopulation`: values → maintained membership (state) -Scenario: keep the full latency population per job, so that p99 and top-10 can be evaluated later. +**What it does.** Keeps every value of a population (not a sketch), so that exact quantiles and top-k can be computed later, and tracks rows entering and leaving. + +**Example.** Keep all latencies per job, so that p99 and top-10 can be computed later (5.5). ```text ( next operator ) @@ -511,14 +531,18 @@ Scenario: keep the full latency population per job, so that p99 and top-10 can b [ Scan t ] closed schema ``` -- Output schema: identical to the child's, all plain. Only `result_kind = State` marks it as maintained state. -- Checks: `population.matches_node(child)`. For `Rows`, the child must be the same closed table `Scan`, the value column must be non-null Float64, and grouping must be `by` with in-range keys. For `CurrentSeries`, it must be a `TimeSeries` scan with the same metric, matchers and grouping labels, under an instant `TimeRange` of `lookback_ms` (which may be omitted only for the default 300 s lookback). -- Coverage: **none**. Maintained membership is not combined by `SummaryMerge`. If maintained populations are later materialized per pane, they derive coverage the same way as `SummaryAgg`. -- Boundary: the output is state because it must also track membership changes; downstream operators can only read it through `EvaluatePopulation`. +| Question | Answer | +|---|---| +| What comes out? | the same columns as the input, all plain. Only the result kind `State` marks it as maintained | +| When is it rejected? | the population description does not match the child. For table rows: the child must be that same table `Scan`, the value column a non-null Float64, and the grouping valid. For PromQL series: a scan of the same metric, labels and grouping, under an instant `TimeRange` | +| What is its coverage? | none. Maintained populations are not merged today | +| State or value? | state, because it must also track membership changes. It is read only through `EvaluatePopulation` | ### 5.5 `EvaluatePopulation`: maintained membership → value -Scenario: p99 and the top-10 latencies by job, both from the one population in 5.4. +**What it does.** Computes an exact statistic from a maintained population. + +**Example.** p99 and the top-10 latencies by job, both from the one population in 5.4. ```text ( next operator ) ( next operator ) @@ -540,24 +564,28 @@ Scenario: p99 and the top-10 latencies by job, both from the one population in 5 [ Scan t ] closed schema ``` -- Output schema: the schema of `Aggregate(by grouping, measure)` over the maintained source. Quantile gives `quantile_` Float64, Sum gives `sum` (value type), Count gives `count` Int64 and Average gives `avg` Float64; `unique_keys = [[0]]`, `closed`. `TopK { k }` instead returns the source schema unchanged (the selected rows). -- Checks: the child is a `MaintainPopulation` node whose `supports(evaluation)` holds: `quantiles` must be set for `Quantile`, and `k <= max_k` for `TopK`. -- Coverage: **none**. -- Boundary: maintained membership is read as a value; the result kind is the source's (`Relation`). +| Question | Answer | +|---|---| +| What comes out? | the same shape as an ordinary `Aggregate` by the grouping: `quantile_0_99` Float64 here, or `sum`, `count`, `avg`. Top-k instead returns the selected rows | +| When is it rejected? | the population was not set up for the question: `quantiles` must be on for a quantile, and `k` must be at most `max_k` for top-k | +| What is its coverage? | none | +| State or value? | value | ### 5.6 `SummaryMerge`: state × N → state (merge and rollup) -Current state: +**What it does.** Combines several states of the same kind into one. With `group_by` (planned) it can also make the grouping coarser (rollup). -- **On `main` (since #560):** `SummaryMerge { children }` is implemented. `validate_inputs()` accepts it when there is at least one child, every child is `State` with exactly one state field, and all children have identical schemas. The output schema is the children's schema. -- **#646 (open):** adds the coverage check. `OperatorNode::new` and `validate_structure` also require equal `definition`s and disjoint selections, and `coverage()` returns the merged coverage. -- **Planned:** `group_by`, so one operator does both merge and rollup: +**Status.** + +- **On `main` (since #560):** implemented. All children must be state with exactly one state field and identical schemas. +- **#646 (open):** also requires equal `definition`s and selections that do not overlap, and computes the merged coverage. +- **Planned:** a `group_by` field, so one operator does both merge and rollup: ```rust SummaryMerge { children: Vec, group_by: Reduction } ``` -**Scenario A, time panes.** Two one-minute KLL panes of PromQL `quantile_over_time(0.99, m[2m])` merge into the two-minute state. Each pane reads `TimeRange(1m)` over `TimeShift(s)` over the scan, as Stage 2 builds them. Both panes share one `Scan`. +**Example A: time panes.** PromQL `quantile_over_time(0.99, m[2m])` built from two one-minute panes, the way Stage 2 builds them. Both panes read the same `Scan`. ```text ( next operator ) @@ -586,7 +614,7 @@ SummaryMerge { children: Vec, group_by: Reduction } [ Scan m ] ``` -Both panes have the same `definition` (`SummaryAgg` over `Scan m`), and their selections are adjacent, so they merge into one range: +Both panes have the same `definition` (`SummaryAgg` over `Scan m`), and their time ranges touch, so the merge covers one continuous range: ```text time −2m −1m 0 @@ -595,7 +623,7 @@ pane 0 (─────────────] merge (───────────────────────────] ``` -**Scenario B, populations.** `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1) merge: +**Example B: regions.** The US and EU states of 5.1 merge into one state for both regions: ```text ( next operator ) @@ -620,7 +648,7 @@ merge (────────────────────── [ Scan t ] ``` -**Scenario C, rollup (planned).** One `KLL(latency) by[region, job]` state merged with `group_by = by[job]`. Each job's output state is the merge of that job's per-region states. The output's `definition` is the same `SummaryAgg` with `reduction = by[job]`, and `selection` is unchanged. The same `by[region, job]` state also answers p99 per region and job directly, so it feeds two consumers. +**Example C: rollup (planned).** A `by[region, job]` state rolled up to `by[job]`: each job's state is the merge of its per-region states. The same `by[region, job]` state also answers p99 per region and job directly. ```text ( next operator ) ( next operator ) @@ -666,45 +694,43 @@ input groups output groups | Merge | Allowed? | Why | Result's selection | |---|---|---|---| -| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (adjacent ranges join) | +| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (touching ranges join) | | `A + C` | ✗ | `(−60s, −30s]` is in both, so those rows would be counted twice | | | `A + A` | ✗ | every row is in both | | | `A + D` | ✗ | different definitions: `D` only has rows with `value * 2 > 10` | | | `A` + a KLL with `k = 400` | ✗ | different definitions (parameters) | | | `(A + B)` + a state over `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | -**How inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): +**Whether inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): | Rule | Families | Example | |---|---|---| | **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | | **may overlap** | HLL, exact `Min`/`Max`, distinct sets | HLL over `A`'s and `C`'s rows ✓: a value seen twice is still one distinct value; the result covers `(−90s, 0]` | -| **right inside left** | subtraction | §5.7 | +| **right inside left** | subtraction | 5.7 | + +| Question | Answer | +|---|---| +| What comes out? | the children's schema; with `group_by`, only the remaining group columns | +| When is it rejected? | no children; a child is not state; the definitions differ (different column, parameters, filters or source); the selections overlap where the family does not allow it; `group_by` is not a subset of the children's grouping | +| What is its coverage? | the shared `definition` (with the new grouping), and the union of the children's selections. Touching ranges join; gaps stay as separate pieces | +| State or value? | state in, state out | -**Coverage outside `SummaryMerge`.** The planner also uses coverage to read or reuse a state: +**Other uses of coverage.** The planner also uses coverage to read or reuse a state without merging: - **Slice:** from a `by[region, job]` state, a query for `region = 'us'` by job reads only the `us` groups ✓. A query for `value < 50` ✗: `value` is not a grouping column, and a sketch cannot be filtered after it is built. - **Reuse:** a stored state answers a query when the definitions are equal and the query's rows are all in the state, with any difference covered by a slice. A stored `by[region, job]` state over `(−5m, 0]` answers p99 by job over the last 5 minutes for `region = 'us'` ✓, but not over the last 1 minute ✗: time is not a grouping column, so minutes 2–5 cannot be taken out. -- Output schema: the children's schema with the group key fields reduced to `group_by`. -- Checks: - - at least one child, every child is `State` with exactly one state field; - - all children have equal `definition`s (§4.2.2), so family, parameters, `input`, `C` and grouping match. Merging k=200 with k=300, KLL over `latency` with KLL over `size`, or states over different sources fails; - - `group_by` ⊆ the children's `G`, and the family merges; - - the children's selections relate as the family requires: disjoint for KLL, so pane 0 with pane 0 is rejected; overlap is allowed for HLL. -- Coverage: **derived**: the shared `definition` with `group_by`, and the union of the children's selections. Adjacent intervals join; gaps stay as separate boxes. Nested merges work because a child merge has coverage like any other summary node. -- Boundary: state in, state out. No value is produced until a readout. - ### 5.7 Reserved operators (not implemented) -These variants exist so that plans can name them, but `output_schema()`/`validate_inputs()` return `UNIMPLEMENTED_ASAP_OP`, so no node can be built. `output_kind()` already returns `State` for each of them. The intended edge shapes below follow from their fields; none of them is implemented. +These operators exist in the code so that plans can name them, but the planner cannot build them yet. Their intended shapes: -| Operator | Fields | Intended edge shape | +| Operator | Inputs | What it would do | |---|---|---| -| `SummarySubtract` | `left, right` | State × State → State: remove one window's contribution, e.g. [0,10) − [0,5). Same `definition`; the right selection must lie inside the left (§4.2) | -| `SummaryDelete` | `summary_input, key: ColumnId` | State → State with the entries for `key` removed | -| `SummaryJoin` | `outer, inner, key, family` | State × State → State typed `family` (`produced_state()` returns it), e.g. join-size estimation | -| `Extension` | `child, name` | deployment-named state operator | +| `SummarySubtract` | `left`, `right` | remove one state from another, e.g. a window minus its oldest part. Same `definition`; the right selection must be inside the left | +| `SummaryDelete` | `summary_input`, `key` | remove the entries for one key | +| `SummaryJoin` | `outer`, `inner`, `key`, `family` | combine two states into a new one, e.g. to estimate a join's size | +| `Extension` | `child`, `name` | a state operator named by the deployment | `SummarySubtract`, for a family that allows it (e.g. an exact `Sum`): From 295617409dfdd37deb7837f240f67215e025bbe8 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:43:01 +0000 Subject: [PATCH 44/59] =?UTF-8?q?docs:=20explain=20every=20struct=20and=20?= =?UTF-8?q?variant=20field=20in=20=C2=A76?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 272 +++++++++++++++--- 1 file changed, 227 insertions(+), 45 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 20e6f4818..7b3c33e57 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -743,18 +743,28 @@ result (──────────────── ## 6. Key code interfaces -`OperatorNode` and coverage are in §6.5. Bodies and serde/derive attributes are elided below. +Every struct field and enum-variant field below has a comment saying what it holds. Function bodies and serde/derive attributes are left out. ### 6.1 Schema and field types (`crates/types/src/pre_asap/schema.rs`) ```rust +/// Position of a column in one schema (0-based). Local to that schema: +/// the same column can have a different `ColumnId` after a projection or join. pub type ColumnId = usize; +/// The columns that flow along one DAG edge. Metadata only: it holds no data. pub struct Schema { + /// The columns, in order. pub fields: Vec, - pub time_index: Option, // must point at a plain Timestamp field + /// The column that holds each row's timestamp, if any. Must point at a + /// plain `Timestamp` field. PromQL inputs always have one. + pub time_index: Option, + /// Sets of columns whose values together identify at most one row, + /// e.g. `[[0]]` when column 0 is unique (one row per `job`). pub unique_keys: Vec>, - pub closed: bool, // true = fields enumerate every column + /// `true`: `fields` lists every column. `false`: more columns may exist + /// that are not listed (e.g. PromQL labels not yet known). + pub closed: bool, } impl Schema { pub fn new(fields: Vec) -> Self; @@ -765,10 +775,19 @@ impl Schema { pub fn column_id_qualified(&self, table: &str, name: &str) -> Option; } +/// One column of a `Schema`. `T` is `FieldDataType` on DAG edges, and plain +/// `DataType` for fields nested inside a `List` or `Struct`. pub struct Field { + /// Column name as the producer outputs it: a SQL column name, or a PromQL + /// label name, `value` or `timestamp`. pub name: String, + /// The column's type: a plain value, or summary state (`FieldDataType`). pub dtype: T, + /// Whether the column may contain NULL. PromQL value columns are never NULL. pub nullable: bool, + /// Table or alias the column comes from (`t` in `t.col`), so two columns + /// with the same name from a join stay apart. `None` for PromQL labels and + /// unqualified columns. pub table: Option, } impl Field { @@ -777,20 +796,39 @@ impl Field { pub fn is_plain(&self) -> bool; } -/// A column's type: a plain value, or summary state of one family. +/// A column's type: a plain value, or summary state of one family (§3). pub enum FieldDataType { + /// A readable value of this type. Plain(DataType), + /// Exact accumulator state: which accumulator, and its parameters. ExactAggregate(ExactKind, ExactParams), + /// Sketch state: the chosen sketch (category, algorithm, parameters), and + /// whether each group has its own sketch or all groups share one. Sketch(SketchKind, GroupingStrategy), + /// Sample state: the sampling method, and its parameters. Sample(SamplingKind, SamplingParams), + /// Wavelet state: the transform, and its parameters. Wavelet(WaveletKind, WaveletParams), + /// Statistical-model state: the model kind, and its parameters. StatModel(StatModelKind, StatModelParams), } +/// Types of plain values. pub enum DataType { - Null, Int64, Float64, Utf8, Bool, Timestamp, Interval, Date, + Null, // only NULL values + Int64, // 64-bit integer + Float64, // 64-bit float + Utf8, // string + Bool, // boolean + Timestamp, // point in time + Interval, // a duration; used for literals, never as a column type + Date, // calendar date without time of day + /// A list. `element`: name, type and nullability of each element. List { element: Box> }, + /// A record. `fields`: its named fields, in order. Struct { fields: Vec> }, + /// A SQL map. `key`: key type (keys are never NULL). `value`: value type. + /// `value_nullable`: whether values may be NULL. Map { key: Box, value: Box, value_nullable: bool }, } ``` @@ -798,10 +836,20 @@ pub enum DataType { ### 6.2 State-family parameters (`crates/types/src/post_asap/sketch.rs`) ```rust +/// Which exact accumulator. None of them has parameters, so `ExactParams` +/// mirrors `ExactKind` one to one. pub enum ExactKind { Sum, Count, Min, Max, Increase, Rate, IRate } -pub enum ExactParams { Sum, Count, Min, Max, Increase, Rate, IRate } // no knobs; mirrors kind - -pub struct SketchKind { category: SketchCategory, algorithm: SketchAlgorithm, params: SketchParams } +pub enum ExactParams { Sum, Count, Min, Max, Increase, Rate, IRate } + +/// A chosen sketch. Built only through `new`, so the three fields always agree. +pub struct SketchKind { + /// What kind of question the sketch answers; derived from `algorithm`. + category: SketchCategory, + /// Which sketch algorithm. + algorithm: SketchAlgorithm, + /// That algorithm's size parameters; must match `algorithm`. + params: SketchParams, +} impl SketchKind { /// The only constructor; classifies the category and panics on mismatched params. pub fn new(algorithm: SketchAlgorithm, params: SketchParams) -> Self; @@ -813,80 +861,147 @@ pub enum SketchCategory { Universal, Quantile, Cardinality, Frequency, TopK } // Universal: UnivMon | Quantile: Kll, DDSketch | Cardinality: Hll, Theta, Kmv // Frequency: Cms, CountSketch | TopK: CmsWithHeap, CountSketchWithHeap pub enum SketchAlgorithm { UnivMon, Kll, Cms, Hll, DDSketch, CmsWithHeap, Kmv, Theta, CountSketch, CountSketchWithHeap } + +/// Size parameters of each algorithm. Larger values: more memory, less error. pub enum SketchParams { + /// `layers`: number of sampling levels. Each level has a Count Sketch of + /// `sketch_rows` hash rows × `sketch_cols` counters, and a heap of the + /// `heap_size` heaviest items. UnivMon { heap_size: u32, sketch_rows: u32, sketch_cols: u32, layers: u8 }, + /// `k`: compactor capacity; error shrinks roughly as 1/k. Kll { k: u32 }, + /// `width`: counters per row. `depth`: number of hash rows. Cms { width: u32, depth: u32 }, + /// `precision`: log2 of the number of registers. Hll { precision: u8 }, + /// `alpha`: relative error of each quantile. DDSketch { alpha: f64 }, + /// Count-Min `width` × `depth`, plus a heap of the `heap_size` heaviest items. CmsWithHeap { width: u32, depth: u32, heap_size: u32 }, + /// `k`: number of smallest hash values kept. Kmv { k: u32 }, + /// `k`: number of hash values kept (nominal entries). Theta { k: u32 }, + /// `width`: counters per row. `depth`: number of hash rows. CountSketch { width: u32, depth: u32 }, + /// Count Sketch `width` × `depth`, plus a heap of the `heap_size` heaviest items. CountSketchWithHeap { width: u32, depth: u32, heap_size: u32 }, } -/// How grouped state is instantiated across `by` subpopulations. Orthogonal to family. +/// How a grouped sketch is laid out across the `by` groups. Independent of the family. pub enum GroupingStrategy { - PerSubpopulationInstance, // Default + /// One separate sketch per group. The default. + PerSubpopulationInstance, + /// One shared structure for all groups (Hydra). `kind`: which shared + /// structure. `params`: its sizes. SharedMultiSubpopulation { kind: HydraKind, params: HydraParams }, } pub enum HydraKind { HydraKll /* experimental, no error bound */, HydraCms, HydraCountSketch } +/// Sizes of the shared structure. pub enum HydraParams { + /// `k`: KLL `k` each group sees. `shared_buckets`: size of the one structure + /// shared by all groups. HydraKll { k: u32, shared_buckets: u32 }, + /// `width` × `depth`: the sketch each group sees. `shared_rows` × + /// `shared_columns`: the one physical grid all groups hash into. HydraCms { width: u32, depth: u32, shared_rows: u32, shared_columns: u32 }, + /// Same fields as `HydraCms`, for Count Sketch. HydraCountSketch { width: u32, depth: u32, shared_rows: u32, shared_columns: u32 }, } pub fn hydra_kind_for(a: &SketchAlgorithm) -> Option; // Cms, CountSketch only +/// `size`: number of rows kept in the reservoir. pub enum SamplingKind { Reservoir } pub enum SamplingParams { Reservoir { size: u32 } } +/// `coefficients`: number of wavelet coefficients kept. pub enum WaveletKind { Haar } pub enum WaveletParams { Haar { coefficients: u32 } } +/// `family`: name of the parametric distribution, e.g. a normal distribution. pub enum StatModelKind { Parametric } pub enum StatModelParams { Parametric { family: String } } ``` ### 6.3 Update input and readouts (`post_asap/sketch.rs`, `post_asap/maintained_population.rs`) ```rust -/// One state update: `item` keys the update for keyed families; `weight` is applied to state. +/// What one input row adds to the state. pub struct SummaryUpdate { + /// The key the row is counted under, for keyed families (e.g. the item + /// in a Count-Min or HLL). `None` for unkeyed families such as KLL. pub item: Option, + /// The value added to the state: the observed value for a KLL, or the + /// count to add for a Count-Min (often `Constant(1.0)`). pub weight: SummaryInputExpr, - pub weight_domain: WeightDomain, // serde default: UnknownOrSigned + /// Whether `weight` is proven never negative. Some algorithms need that. + /// Missing proof is never treated as non-negative. + pub weight_domain: WeightDomain, } impl SummaryUpdate { pub fn column(c: ColumnRef) -> Self; } // item None, UnknownOrSigned pub enum WeightDomain { - UnknownOrSigned, // Default; never assumed non-negative + /// Nothing is known; the weight may be negative. The default. + UnknownOrSigned, + /// The weight is never negative. `proof`: why. NonNegative { proof: NonNegativeWeightProof }, } -pub enum NonNegativeWeightProof { UnitCount, ResetAwareCounterDerivative } +pub enum NonNegativeWeightProof { + UnitCount, // every row adds 1 + ResetAwareCounterDerivative, // PromQL increase/rate over counters with reset correction +} +/// An expression that computes an item or a weight from the input row. pub enum SummaryInputExpr { - Constant(f64), Column(ColumnRef), Tuple(Vec), EntityIdentity(EntityIdentity), + Constant(f64), // the same number for every row + Column(ColumnRef), // the value of one column + Tuple(Vec), // several values combined into one item + EntityIdentity(EntityIdentity), // the identity of the row's series +} +pub enum EntityIdentity { + /// A PromQL series identified by its labels. `excluding`: labels left out. + PromqlLabelSet { excluding: Vec }, } -pub enum EntityIdentity { PromqlLabelSet { excluding: Vec } } -/// Readout of sketch state, carried by SummaryEstimate. +/// What `SummaryEstimate` reads out of a sketch. pub enum SketchStatistic { - FrequencyL2, FrequencyEntropy, + FrequencyL2, // sqrt of the sum of squared item frequencies + FrequencyEntropy, // entropy of the item frequencies, in bits + /// `q`: the quantile rank, in (0, 1]. Quantile { q: f64 }, + /// An estimated count. `key`: the column being counted. `value`: the item + /// to look up (e.g. `"checkout"`), or `None` for the total count. PointCount { key: ColumnRef, value: Option }, - Cardinality, + Cardinality, // number of distinct items + /// `k`: how many top items to return. TopK { k: usize }, } -/// Readout of a maintained population, carried by EvaluatePopulation. -pub enum PopulationStatistic { Quantile { q: f64 }, TopK { k: usize }, Sum, Count, Average } +/// What `EvaluatePopulation` computes from a maintained population. +pub enum PopulationStatistic { + Quantile { q: f64 }, // `q`: the quantile rank + TopK { k: usize }, // `k`: how many top rows to return + Sum, Count, Average, +} +/// A population kept in full (§5.4). pub struct MaintainedPopulation { + /// Which rows or series the population contains. pub input: PopulationInput, - pub max_k: usize, // largest TopK it supports - pub quantiles: bool, // whether Quantile is supported + /// The largest `k` a `TopK` read may ask for. + pub max_k: usize, + /// Whether `Quantile` reads are supported. + pub quantiles: bool, } pub enum PopulationInput { - CurrentSeries(CurrentSeriesInput), // metric, matchers, grouping, without, lookback_ms + /// The current value of each PromQL series (see `CurrentSeriesInput`). + CurrentSeries(CurrentSeriesInput), + /// Rows of a table. `input`: the table `Scan` node. `value_column`: the + /// column whose values are kept. `grouping`: the `by` columns. Rows { input: Rc, value_column: usize, grouping: GroupKeys }, } +pub struct CurrentSeriesInput { + pub metric: String, // metric name + pub matchers: Vec, // label matchers, e.g. job="api" + pub grouping: Vec, // labels to group by + pub without: bool, // true: group by all labels except `grouping` + pub lookback_ms: u64, // how long a series stays current without new samples +} ``` -A finalized value's accuracy statement is `ResultGuarantee { metric, bound, failure_probability, provenance }` (`post_asap/guarantee.rs`). It is attached to readout and finalized nodes, never to raw state. +A finalized value's accuracy statement is `ResultGuarantee` (`post_asap/guarantee.rs`): `metric` (which error is measured, e.g. rank error), `bound` (the error bound), `failure_probability` (the chance the bound does not hold) and `provenance` (which estimates the bound came from). It is attached to readout and finalized nodes, never to raw state. ### 6.4 ASAP operators (`crates/types/src/ir/asap.rs`) @@ -896,24 +1011,66 @@ pub const UNIMPLEMENTED_ASAP_OP: &str = pub enum ASAPOp { SummaryAgg { + /// The input rows. child: Rc, - family: FieldDataType, // never Plain + /// The summary type of the output `state` field. Never `Plain`. + family: FieldDataType, + /// What each input row adds to the state (item and weight). input: SummaryUpdate, - reduction: Reduction, // Reduce(GroupKeys) | PerEntity + /// The grouping: `Reduce(by columns)`, or `PerEntity` for one state per + /// input series without grouping. + reduction: Reduction, + /// Whether each group gets its own sketch or all groups share one. grouping: GroupingStrategy, - filter: Option, // serde default None + /// Rows to include, applied before updating the state. `None`: all rows. + filter: Option, + }, + SummaryEstimate { + /// The node that produces the sketch state. + summary_input: Rc, + /// What to read out of it. + query: SketchStatistic, + }, + FinalizeExactAccumulator { + /// The node that produces the exact accumulator state. + child: Rc, + }, + MaintainPopulation { + /// The input rows; must match `population.input`. + child: Rc, + /// What population to keep and which reads it supports. + population: MaintainedPopulation, + }, + EvaluatePopulation { + /// The `MaintainPopulation` node. + child: Rc, + /// What to compute from it. + evaluation: PopulationStatistic, }, - SummaryEstimate { summary_input: Rc, query: SketchStatistic }, - FinalizeExactAccumulator { child: Rc }, - MaintainPopulation { child: Rc, population: MaintainedPopulation }, - EvaluatePopulation { child: Rc, evaluation: PopulationStatistic }, // Implemented since #560 (identical child schemas). - SummaryMerge { children: Vec> }, + SummaryMerge { + /// The states to merge; all have the same schema. + children: Vec>, + }, // Reserved: migrated but unimplemented. - SummarySubtract { left: Rc, right: Rc }, - SummaryDelete { summary_input: Rc, key: ColumnId }, - SummaryJoin { outer: Rc, inner: Rc, key: ColumnId, family: FieldDataType }, - Extension { child: Rc, name: String }, + SummarySubtract { + left: Rc, // the state to subtract from + right: Rc, // the state to remove from `left` + }, + SummaryDelete { + summary_input: Rc, // the state + key: ColumnId, // the key column whose entries are removed + }, + SummaryJoin { + outer: Rc, // one input state + inner: Rc, // the other input state + key: ColumnId, // the join key column + family: FieldDataType, // the summary type of the result + }, + Extension { + child: Rc, // the input + name: String, // the deployment-defined operator name + }, } impl ASAPOp { @@ -935,15 +1092,24 @@ impl ASAPOp { The code for §4. ```rust +/// One node of the DAG. Immutable and shared through `Rc`. pub struct OperatorNode { + /// What the node does: an ordinary operator or an ASAP operator. pub operator: Operator, + /// What kind of output it has: `Relation`, `InstantVector`, + /// `RangeVector`, or `State` (summary state). Derived from `operator`. pub result_kind: OperatorResultKind, + /// The output columns. Derived from `operator` and its children. pub schema: Schema, + /// The accuracy statement of the output, once known. `None` does not + /// mean exact. pub guarantee: Option, + /// When the node runs: `IngestionTime` or `QueryTime`. `None` until + /// planning assigns it. pub timing: Option, - /// Cache for `coverage()`. Lazily filled, never serialized, ignored by - /// equality, emptied on clone. Not a source of truth: coverage is always - /// re-derivable. + /// Cache for `coverage()`. Filled on first use, never serialized, + /// ignored by equality, emptied on clone. Not a source of truth: + /// coverage can always be derived again. coverage_cache: CoverageCache, } @@ -953,25 +1119,41 @@ impl OperatorNode { } pub struct SummaryCoverage { - /// The `SummaryAgg` (or rolled-up equivalent) with the selection removed. + /// What the state computes: the `SummaryAgg` and its sub-DAG, with the + /// conditions that went into `selection` taken out (§4.2.2). pub definition: Rc, - /// Union of boxes over the output rows of the definition's computation. + /// Which rows went in: a union of boxes. A row is in the state if it is + /// in at least one box. pub selection: Vec, } +/// One box: a row is in it when it meets every constraint. pub struct SelectionBox { - pub columns: BTreeMap, // missing column = unrestricted - pub relative_time: Option<(Bound, Bound)>, // ms from evaluation; None = unrestricted + /// One constraint per restricted column. A column not in the map is + /// unrestricted. + pub columns: BTreeMap, + /// The time window relative to the evaluation time, in ms, e.g. + /// `(Excluded(-120000), Included(-60000))` for `(−2m, −1m]`. + /// `None`: no time restriction. + pub relative_time: Option<(Bound, Bound)>, } +/// Names a column across nodes, independent of its position. pub struct ColumnIdentity { + /// The table or alias the column comes from; `None` if unqualified. pub table: Option, + /// The column name. pub name: String, } +/// A constraint on one column. pub enum Constraint { - In(Vec), // ScalarValue has no total order (Float64) + /// The value is one of these. + In(Vec), + /// The value is none of these. NotIn(Vec), + /// The value lies between `lower` and `upper`. Each end is `Included`, + /// `Excluded` or `Unbounded`. Interval { lower: Bound, upper: Bound }, // HashPartition { columns, of, index }: added with its first producer. } From 75f508f1ab23a8e798fdbb37fab2b5b91c4c5ea5 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:43:14 +0000 Subject: [PATCH 45/59] =?UTF-8?q?docs:=20label=20=C2=A75=20operator=20desc?= =?UTF-8?q?riptions=20as=20operator=20definitions?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 7b3c33e57..60b9eaf12 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -410,7 +410,7 @@ All examples read a table, so values are `Relation`. For PromQL series they woul ### 5.1 `SummaryAgg`: values → state -**What it does.** Turns rows into summary state: one state per group. +**Operator definition.** Turns rows into summary state: one state per group. **Example.** p99 latency by job, from a KLL sketch with `k = 200`, over table `t`, using only US rows with latency under 10 s. @@ -445,7 +445,7 @@ coverage() of the SummaryAgg ### 5.2 `SummaryEstimate`: sketch state → value -**What it does.** Reads a number out of a sketch, for example a quantile or a count. +**Operator definition.** Reads a number out of a sketch, for example a quantile or a count. **Example.** Read p99 and p50 from the state in 5.1. One state feeds both readouts. @@ -480,7 +480,7 @@ coverage() of the SummaryAgg ### 5.3 `FinalizeExactAccumulator`: exact state → value -**What it does.** Turns an exact accumulator (sum, count, min, max, rate, …) into its final value. +**Operator definition.** Turns an exact accumulator (sum, count, min, max, rate, …) into its final value. **Example.** Total bytes by host, with an exact `Sum` accumulator. @@ -515,7 +515,7 @@ coverage() of the SummaryAgg ### 5.4 `MaintainPopulation`: values → maintained membership (state) -**What it does.** Keeps every value of a population (not a sketch), so that exact quantiles and top-k can be computed later, and tracks rows entering and leaving. +**Operator definition.** Keeps every value of a population (not a sketch), so that exact quantiles and top-k can be computed later, and tracks rows entering and leaving. **Example.** Keep all latencies per job, so that p99 and top-10 can be computed later (5.5). @@ -540,7 +540,7 @@ coverage() of the SummaryAgg ### 5.5 `EvaluatePopulation`: maintained membership → value -**What it does.** Computes an exact statistic from a maintained population. +**Operator definition.** Computes an exact statistic from a maintained population. **Example.** p99 and the top-10 latencies by job, both from the one population in 5.4. @@ -573,7 +573,7 @@ coverage() of the SummaryAgg ### 5.6 `SummaryMerge`: state × N → state (merge and rollup) -**What it does.** Combines several states of the same kind into one. With `group_by` (planned) it can also make the grouping coarser (rollup). +**Operator definition.** Combines several states of the same kind into one. With `group_by` (planned) it can also make the grouping coarser (rollup). **Status.** From b00fb0a34275edd16a6778fe46e7dd6e92283699 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:58:29 +0000 Subject: [PATCH 46/59] =?UTF-8?q?docs:=20explain=20every=20function=20in?= =?UTF-8?q?=20=C2=A76?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 59 ++++++++++++++++--- 1 file changed, 51 insertions(+), 8 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 60b9eaf12..84f45bec5 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -767,11 +767,21 @@ pub struct Schema { pub closed: bool, } impl Schema { + /// A schema with only these fields: no time column, no unique key, not + /// closed. Used for a table `Scan` without key metadata. pub fn new(fields: Vec) -> Self; + /// A schema with a time column and unique keys, not closed. Used for + /// PromQL inputs, whose rows are unique per (time, label set). pub fn with_time_index(fields: Vec, time_index: ColumnId, unique_keys: Vec>) -> Self; - pub fn lifted(fields: Vec, time_index: Option) -> Self; // closed = true + /// A closed schema with an optional time column and no unique key: the + /// shape of a summary operator's output. + pub fn lifted(fields: Vec, time_index: Option) -> Self; + /// Whether every field holds a plain value (no summary state). pub fn is_all_plain(&self) -> bool; + /// Position of the first field with this name, or `None`. pub fn column_id(&self, name: &str) -> Option; + /// Position of the field with this table and name, or `None`. Tells apart + /// `a.k` and `b.k` after a join. pub fn column_id_qualified(&self, table: &str, name: &str) -> Option; } @@ -791,8 +801,11 @@ pub struct Field { pub table: Option, } impl Field { + /// An unqualified field holding a plain value of type `dtype`. pub fn plain(name: impl Into, dtype: DataType, nullable: bool) -> Self; + /// The value type if the field is plain; `None` for summary state. pub fn plain_dtype(&self) -> Option<&DataType>; + /// Whether the field holds a plain value (not summary state). pub fn is_plain(&self) -> bool; } @@ -851,10 +864,14 @@ pub struct SketchKind { params: SketchParams, } impl SketchKind { - /// The only constructor; classifies the category and panics on mismatched params. + /// The only constructor. Derives the category from the algorithm, and + /// panics if `params` belong to a different algorithm. pub fn new(algorithm: SketchAlgorithm, params: SketchParams) -> Self; + /// What kind of question the sketch answers. pub fn category(&self) -> SketchCategory; + /// The sketch algorithm. pub fn algorithm(&self) -> &SketchAlgorithm; + /// The algorithm's size parameters. pub fn params(&self) -> &SketchParams; } pub enum SketchCategory { Universal, Quantile, Cardinality, Frequency, TopK } @@ -908,7 +925,10 @@ pub enum HydraParams { /// Same fields as `HydraCms`, for Count Sketch. HydraCountSketch { width: u32, depth: u32, shared_rows: u32, shared_columns: u32 }, } -pub fn hydra_kind_for(a: &SketchAlgorithm) -> Option; // Cms, CountSketch only +/// The shared (Hydra) version of an algorithm, if one with a known error +/// bound exists: `Cms` → `HydraCms`, `CountSketch` → `HydraCountSketch`; +/// `None` for all others. +pub fn hydra_kind_for(a: &SketchAlgorithm) -> Option; /// `size`: number of rows kept in the reservoir. pub enum SamplingKind { Reservoir } pub enum SamplingParams { Reservoir { size: u32 } } @@ -933,7 +953,11 @@ pub struct SummaryUpdate { /// Missing proof is never treated as non-negative. pub weight_domain: WeightDomain, } -impl SummaryUpdate { pub fn column(c: ColumnRef) -> Self; } // item None, UnknownOrSigned +impl SummaryUpdate { + /// Add the value of column `c` for each row: no item, and the weight is + /// not known to be non-negative. + pub fn column(c: ColumnRef) -> Self; +} pub enum WeightDomain { /// Nothing is known; the weight may be negative. The default. UnknownOrSigned, @@ -1074,15 +1098,27 @@ pub enum ASAPOp { } impl ASAPOp { - pub fn children(&self) -> Vec<&Rc>; // SummaryAgg includes its filter's subquery nodes + /// The input nodes. For `SummaryAgg` this also includes nodes used by + /// subqueries inside its `filter`. + pub fn children(&self) -> Vec<&Rc>; + /// The same operator with each input replaced by `f(input)`. pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Self; + /// The operator's name, e.g. `"SummaryAgg"`, for messages and display. pub fn kind_name(&self) -> &'static str; - /// Subtract, Delete, Join, Extension. + /// Whether the operator is reserved and cannot be built yet: Subtract, + /// Delete, Join, Extension. pub fn is_unimplemented(&self) -> bool; - /// SummaryAgg/SummaryJoin `family`; SummaryMerge: its inputs' state type. + /// The summary type this operator outputs: `family` for SummaryAgg and + /// SummaryJoin, the inputs' state type for SummaryMerge, `None` otherwise. pub fn produced_state(&self) -> Option<&FieldDataType>; + /// The output schema, derived from the operator and its inputs. An error + /// for a reserved operator or an invalid input. pub fn output_schema(&self) -> Result; + /// The output kind: `State` for operators that output state, otherwise + /// the value kind of the input. pub fn output_kind(&self) -> OperatorResultKind; + /// Checks the inputs (§5, "When is it rejected?"). An error if they do + /// not fit the operator. pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>; } ``` @@ -1114,7 +1150,9 @@ pub struct OperatorNode { } impl OperatorNode { - /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`. + /// The node's coverage (§4.2), derived on first use and then cached. + /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`; `None` for every + /// other node. pub fn coverage(&self) -> Option<&SummaryCoverage>; } @@ -1159,6 +1197,11 @@ pub enum Constraint { } impl SummaryCoverage { + /// Computes the coverage of a summary node from its sub-DAG (§4.2.2). + /// Errors: `NotSummary` for a node that is not a `SummaryAgg` or + /// `SummaryMerge`; for a merge, `EmptyMerge` (no inputs), + /// `DefinitionMismatch` (inputs compute different things) or + /// `PossibleOverlap` (inputs may share rows). pub fn derive(node: &OperatorNode) -> Result; } ``` From 4595032e192a8a51d29d4e78cbfc5333ecf9de7b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 20:59:44 +0000 Subject: [PATCH 47/59] docs: describe SketchKind.category as the aggregation intents it answers Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 84f45bec5..4984b8309 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -856,7 +856,10 @@ pub enum ExactParams { Sum, Count, Min, Max, Increase, Rate, IRate } /// A chosen sketch. Built only through `new`, so the three fields always agree. pub struct SketchKind { - /// What kind of question the sketch answers; derived from `algorithm`. + /// Which aggregation intents the sketch answers: `Quantile` (quantiles), + /// `Cardinality` (distinct counts), `Frequency` (item counts), `TopK` + /// (heavy hitters), or `Universal` (frequency moments such as L2 and + /// entropy, plus counts and distinct counts). Derived from `algorithm`. category: SketchCategory, /// Which sketch algorithm. algorithm: SketchAlgorithm, @@ -867,7 +870,7 @@ impl SketchKind { /// The only constructor. Derives the category from the algorithm, and /// panics if `params` belong to a different algorithm. pub fn new(algorithm: SketchAlgorithm, params: SketchParams) -> Self; - /// What kind of question the sketch answers. + /// Which aggregation intents the sketch answers (see the `category` field). pub fn category(&self) -> SketchCategory; /// The sketch algorithm. pub fn algorithm(&self) -> &SketchAlgorithm; From 111bfe907ec9ed8fb6aa5323b8ccd0a8e634397b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:01:44 +0000 Subject: [PATCH 48/59] =?UTF-8?q?docs:=20explain=20what=20=C2=A76.3=20upda?= =?UTF-8?q?te=20input=20and=20readouts=20are=20for,=20with=20examples?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 4984b8309..0efb2efc5 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -943,6 +943,23 @@ pub enum StatModelKind { Parametric } pub enum StatModelParams { Parametric { ### 6.3 Update input and readouts (`post_asap/sketch.rs`, `post_asap/maintained_population.rs`) +This part of the code answers two questions about summary state: + +1. **Update input** (`SummaryUpdate`): when a `SummaryAgg` reads one input row, what does it add to the state? It adds a **weight**, optionally under an **item** key. +2. **Readouts**: once the state is built, what can be asked of it? `SketchStatistic` is what `SummaryEstimate` asks a sketch (§5.2). `MaintainedPopulation` and `PopulationStatistic` are the exact-population counterpart, used by `MaintainPopulation` and `EvaluatePopulation` (§5.4, §5.5). + +**Update input examples** (as the planner builds them in `asap-aware-mapping`): + +| Query | Sketch | `item` | `weight` | `weight_domain` | +|---|---|---|---|---| +| p99 of `latency` | KLL | none: KLL has no keys | `Column(latency)`: the value itself | `UnknownOrSigned` | +| how often each `endpoint` occurs | Count-Min | `Column(endpoint)` | `Constant(1.0)`: each row counts once | `NonNegative(UnitCount)` | +| `topk(5, rate(http_requests_total[5m]))` | Count-Min with heap | `Tuple(label columns)`: one item per series | `Column(value)`: the rate | `NonNegative(ResetAwareCounterDerivative)` | + +`weight_domain` matters because some sketches (e.g. Count-Min) are only accurate when weights are never negative. The planner records why a weight is non-negative; if it cannot prove it, the weight counts as possibly negative. + +**Readout examples:** `Quantile { q: 0.99 }` reads p99 from a KLL; `PointCount { key: endpoint, value: Some("checkout") }` reads how often `checkout` occurred from a Count-Min; `TopK { k: 5 }` reads the 5 heaviest items. + ```rust /// What one input row adds to the state. pub struct SummaryUpdate { From 3cb0ec5723a51dbd47955132fc342f9bf00ee08e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:01:59 +0000 Subject: [PATCH 49/59] =?UTF-8?q?docs:=20rename=20=C2=A76.2=20to=20summary?= =?UTF-8?q?-family=20parameters?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 0efb2efc5..f1e8ff744 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -77,7 +77,7 @@ Based on our requirement, each field should contain the following information. | Algorithm | `SketchAlgorithm` (other families: `ExactKind`, `SamplingKind`, …) | `Kll`, `Cms`, `Hll`, `DDSketch`, … | `Kll` | | Parameters | `SketchParams` (other families: `ExactParams`, `SamplingParams`, …) | per algorithm | `Kll { k: 200 }` | - - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.2](#62-state-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. + - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.2](#62-summary-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. - So a quantile KLL sketch with `k = 200`, one instance per group, has the type `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, PerSubpopulationInstance)`. 2. **What query intent the summarized ASAP Primitive can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: @@ -846,7 +846,7 @@ pub enum DataType { } ``` -### 6.2 State-family parameters (`crates/types/src/post_asap/sketch.rs`) +### 6.2 Summary-family parameters (`crates/types/src/post_asap/sketch.rs`) ```rust /// Which exact accumulator. None of them has parameters, so `ExactParams` From aeaf98357da487ce2c60e2dbe0a178ca016887d0 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:08:40 +0000 Subject: [PATCH 50/59] =?UTF-8?q?docs:=20show=20how=20the=20=C2=A76=20code?= =?UTF-8?q?=20interfaces=20connect?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 30 +++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index f1e8ff744..2fbfc88e0 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -745,6 +745,36 @@ result (──────────────── Every struct field and enum-variant field below has a comment saying what it holds. Function bodies and serde/derive attributes are left out. +**How the subsections connect.** Everything hangs off one DAG node, `OperatorNode` (§6.5): + +```text +OperatorNode §6.5 +├── operator: Operator::ASAP(ASAPOp) §6.4 +│ ├── SummaryAgg +│ │ ├── family: FieldDataType ───────────┐ §6.1 → §6.2 +│ │ └── input: SummaryUpdate │ §6.3 (what each row adds) +│ ├── SummaryEstimate.query: SketchStatistic│ §6.3 (what is read out) +│ ├── MaintainPopulation.population │ §6.3 +│ └── EvaluatePopulation.evaluation │ §6.3 +├── schema: Schema │ §6.1 +│ └── fields[i].dtype: FieldDataType ◀───────┘ §6.1 (the same type: the output +│ └── Sketch(SketchKind, GroupingStrategy) `state` field has type `family`) +│ └── category, algorithm, params §6.2 +├── guarantee: ResultGuarantee §6.3 (accuracy of a readout) +└── coverage() → SummaryCoverage §6.5 + └── definition: OperatorNode (again a node, so the same structure) +``` + +| Subsection | Defines | Used by | +|---|---|---| +| §6.1 Schema and field types | `Schema`, `Field`, `FieldDataType` | every node's `schema`; `SummaryAgg.family` | +| §6.2 Summary-family parameters | `SketchKind`, `SketchParams`, `GroupingStrategy`, exact/sample/wavelet/model params | the payload of each non-plain `FieldDataType` variant | +| §6.3 Update input and readouts | `SummaryUpdate`, `SketchStatistic`, `MaintainedPopulation`, `PopulationStatistic` | fields of the `ASAPOp` variants in §6.4 | +| §6.4 ASAP operators | `ASAPOp` | `OperatorNode.operator` | +| §6.5 Summary coverage | `OperatorNode`, `SummaryCoverage` | the DAG itself; `coverage()` on summary nodes | + +**Example: the p99 latency state of §5.1.** The `SummaryAgg` node (§6.5) holds `ASAPOp::SummaryAgg` (§6.4) with `family = Sketch(SketchKind::new(Kll, Kll { k: 200 }), PerSubpopulationInstance)` (§6.1, §6.2) and `input = SummaryUpdate::column(latency)` (§6.3). Its `schema` (§6.1) is `(job Utf8, state Sketch(KLL k=200))`: the `state` field has the same type as `family`. A `SummaryEstimate` above it holds `query = Quantile { q: 0.99 }` (§6.3), and its node carries the `guarantee` of the readout. `coverage()` on the `SummaryAgg` node returns a `SummaryCoverage` (§6.5). + ### 6.1 Schema and field types (`crates/types/src/pre_asap/schema.rs`) ```rust From dc96bd3b92a073a70c467563171f3e479779968f Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:11:56 +0000 Subject: [PATCH 51/59] =?UTF-8?q?docs:=20order=20=C2=A76=20subsections=20f?= =?UTF-8?q?rom=20the=20DAG=20node=20inward?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 309 +++++++++--------- 1 file changed, 158 insertions(+), 151 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 2fbfc88e0..66dc8dca0 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -65,7 +65,7 @@ Schema represents the **metadata** of information flow along an **edge** between Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. Based on our requirement, each field should contain the following information. -1. **What type of the ASAP Primitive is.** The field's type is a [`FieldDataType`](#61-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: +1. **What type of the ASAP Primitive is.** The field's type is a [`FieldDataType`](#63-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: - **Raw value**: `Plain(DataType)`, e.g., a number or a string. - **Summary state**: described from coarse to fine by four levels: @@ -77,7 +77,7 @@ Based on our requirement, each field should contain the following information. | Algorithm | `SketchAlgorithm` (other families: `ExactKind`, `SamplingKind`, …) | `Kll`, `Cms`, `Hll`, `DDSketch`, … | `Kll` | | Parameters | `SketchParams` (other families: `ExactParams`, `SamplingParams`, …) | per algorithm | `Kll { k: 200 }` | - - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.2](#62-summary-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. + - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.4](#64-summary-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. - So a quantile KLL sketch with `k = 200`, one instance per group, has the type `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, PerSubpopulationInstance)`. 2. **What query intent the summarized ASAP Primitive can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: @@ -93,7 +93,7 @@ Based on our requirement, each field should contain the following information. | `FrequencyL2`, `FrequencyEntropy` | `Sketch`: `UnivMon` | `SummaryEstimate(FrequencyL2 \| FrequencyEntropy)` | | `Sum`, `Count`, `Min`, `Max`, `Rate`, `Increase` (exact) | `ExactAggregate(ExactKind, …)` | `FinalizeExactAccumulator` | - The sketch candidates are `summary_candidates(intent)` in `crates/asap-aware-mapping/src/replacement.rs`; the readouts are `SketchStatistic` ([§6.3](#63-update-input-and-readouts-post_asapsketchrs-post_asapmaintained_populationrs)). + The sketch candidates are `summary_candidates(intent)` in `crates/asap-aware-mapping/src/replacement.rs`; the readouts are `SketchStatistic` ([§6.5](#65-update-input-and-readouts-post_asapsketchrs-post_asapmaintained_populationrs)). - **Time window aggregation intents**: whether states built over smaller windows can answer a larger one. This depends on how the family combines states: - **Merge** (`SummaryMerge`, §5.6): states over disjoint panes combine into the state of their union, e.g. two 1-minute KLL states answer a 2-minute quantile. Requires a mergeable family. @@ -115,7 +115,7 @@ A node in the physical data will represent the data or summary instance, so a no | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2.1 | | How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.2.2 | | What uses it? | merge, rollup, slice, reuse, subtract | §5.6 | -| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.5 | +| Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.1, §6.6 | | What is left out? | the deployment and runtime implementation, e.g. SDS | §4.3 | **Why coverage is not part of the schema.** @@ -745,37 +745,168 @@ result (──────────────── Every struct field and enum-variant field below has a comment saying what it holds. Function bodies and serde/derive attributes are left out. -**How the subsections connect.** Everything hangs off one DAG node, `OperatorNode` (§6.5): +**How the subsections connect.** The subsections follow one DAG node from the outside in: ```text -OperatorNode §6.5 -├── operator: Operator::ASAP(ASAPOp) §6.4 +OperatorNode §6.1 +├── operator: Operator::ASAP(ASAPOp) §6.2 │ ├── SummaryAgg -│ │ ├── family: FieldDataType ───────────┐ §6.1 → §6.2 -│ │ └── input: SummaryUpdate │ §6.3 (what each row adds) -│ ├── SummaryEstimate.query: SketchStatistic│ §6.3 (what is read out) -│ ├── MaintainPopulation.population │ §6.3 -│ └── EvaluatePopulation.evaluation │ §6.3 -├── schema: Schema │ §6.1 -│ └── fields[i].dtype: FieldDataType ◀───────┘ §6.1 (the same type: the output +│ │ ├── family: FieldDataType ───────────┐ §6.3 → §6.4 +│ │ └── input: SummaryUpdate │ §6.5 (what each row adds) +│ ├── SummaryEstimate.query: SketchStatistic│ §6.5 (what is read out) +│ ├── MaintainPopulation.population │ §6.5 +│ └── EvaluatePopulation.evaluation │ §6.5 +├── schema: Schema │ §6.3 +│ └── fields[i].dtype: FieldDataType ◀───────┘ §6.3 (the same type: the output │ └── Sketch(SketchKind, GroupingStrategy) `state` field has type `family`) -│ └── category, algorithm, params §6.2 -├── guarantee: ResultGuarantee §6.3 (accuracy of a readout) -└── coverage() → SummaryCoverage §6.5 +│ └── category, algorithm, params §6.4 +├── guarantee: ResultGuarantee §6.5 (accuracy of a readout) +└── coverage() → SummaryCoverage §6.6 └── definition: OperatorNode (again a node, so the same structure) ``` | Subsection | Defines | Used by | |---|---|---| -| §6.1 Schema and field types | `Schema`, `Field`, `FieldDataType` | every node's `schema`; `SummaryAgg.family` | -| §6.2 Summary-family parameters | `SketchKind`, `SketchParams`, `GroupingStrategy`, exact/sample/wavelet/model params | the payload of each non-plain `FieldDataType` variant | -| §6.3 Update input and readouts | `SummaryUpdate`, `SketchStatistic`, `MaintainedPopulation`, `PopulationStatistic` | fields of the `ASAPOp` variants in §6.4 | -| §6.4 ASAP operators | `ASAPOp` | `OperatorNode.operator` | -| §6.5 Summary coverage | `OperatorNode`, `SummaryCoverage` | the DAG itself; `coverage()` on summary nodes | +| §6.1 DAG node | `OperatorNode` | the DAG itself | +| §6.2 ASAP operators | `ASAPOp` | `OperatorNode.operator` | +| §6.3 Schema and field types | `Schema`, `Field`, `FieldDataType` | every node's `schema`; `SummaryAgg.family` | +| §6.4 Summary-family parameters | `SketchKind`, `SketchParams`, `GroupingStrategy`, exact/sample/wavelet/model params | the payload of each non-plain `FieldDataType` variant | +| §6.5 Update input and readouts | `SummaryUpdate`, `SketchStatistic`, `MaintainedPopulation`, `PopulationStatistic` | fields of the `ASAPOp` variants in §6.2 | +| §6.6 Summary coverage | `SummaryCoverage`, `SelectionBox`, `Constraint` | `OperatorNode::coverage()` on summary nodes | -**Example: the p99 latency state of §5.1.** The `SummaryAgg` node (§6.5) holds `ASAPOp::SummaryAgg` (§6.4) with `family = Sketch(SketchKind::new(Kll, Kll { k: 200 }), PerSubpopulationInstance)` (§6.1, §6.2) and `input = SummaryUpdate::column(latency)` (§6.3). Its `schema` (§6.1) is `(job Utf8, state Sketch(KLL k=200))`: the `state` field has the same type as `family`. A `SummaryEstimate` above it holds `query = Quantile { q: 0.99 }` (§6.3), and its node carries the `guarantee` of the readout. `coverage()` on the `SummaryAgg` node returns a `SummaryCoverage` (§6.5). +**Example: the p99 latency state of §5.1.** The `SummaryAgg` node (§6.1) holds `ASAPOp::SummaryAgg` (§6.2) with `family = Sketch(SketchKind::new(Kll, Kll { k: 200 }), PerSubpopulationInstance)` (§6.3, §6.4) and `input = SummaryUpdate::column(latency)` (§6.5). Its `schema` (§6.3) is `(job Utf8, state Sketch(KLL k=200))`: the `state` field has the same type as `family`. A `SummaryEstimate` above it holds `query = Quantile { q: 0.99 }` (§6.5), and its node carries the `guarantee` of the readout. `coverage()` on the `SummaryAgg` node returns a `SummaryCoverage` (§6.6). -### 6.1 Schema and field types (`crates/types/src/pre_asap/schema.rs`) +### 6.1 DAG node (`crates/types/src/ir/node.rs`) + +Every operator in a DAG is wrapped in an `OperatorNode`. `coverage()` is explained in §6.6. + +```rust +/// One node of the DAG. Immutable and shared through `Rc`. +pub struct OperatorNode { + /// What the node does: an ordinary operator or an ASAP operator. + pub operator: Operator, + /// What kind of output it has: `Relation`, `InstantVector`, + /// `RangeVector`, or `State` (summary state). Derived from `operator`. + pub result_kind: OperatorResultKind, + /// The output columns. Derived from `operator` and its children. + pub schema: Schema, + /// The accuracy statement of the output, once known. `None` does not + /// mean exact. + pub guarantee: Option, + /// When the node runs: `IngestionTime` or `QueryTime`. `None` until + /// planning assigns it. + pub timing: Option, + /// Cache for `coverage()`. Filled on first use, never serialized, + /// ignored by equality, emptied on clone. Not a source of truth: + /// coverage can always be derived again. + coverage_cache: CoverageCache, +} + +impl OperatorNode { + /// The node's coverage (§4.2), derived on first use and then cached. + /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`; `None` for every + /// other node. + pub fn coverage(&self) -> Option<&SummaryCoverage>; +} +``` + +### 6.2 ASAP operators (`crates/types/src/ir/asap.rs`) + +```rust +pub const UNIMPLEMENTED_ASAP_OP: &str = + "this ASAP operator is reserved: schema, accuracy, timing and export are not implemented"; + +pub enum ASAPOp { + SummaryAgg { + /// The input rows. + child: Rc, + /// The summary type of the output `state` field. Never `Plain`. + family: FieldDataType, + /// What each input row adds to the state (item and weight). + input: SummaryUpdate, + /// The grouping: `Reduce(by columns)`, or `PerEntity` for one state per + /// input series without grouping. + reduction: Reduction, + /// Whether each group gets its own sketch or all groups share one. + grouping: GroupingStrategy, + /// Rows to include, applied before updating the state. `None`: all rows. + filter: Option, + }, + SummaryEstimate { + /// The node that produces the sketch state. + summary_input: Rc, + /// What to read out of it. + query: SketchStatistic, + }, + FinalizeExactAccumulator { + /// The node that produces the exact accumulator state. + child: Rc, + }, + MaintainPopulation { + /// The input rows; must match `population.input`. + child: Rc, + /// What population to keep and which reads it supports. + population: MaintainedPopulation, + }, + EvaluatePopulation { + /// The `MaintainPopulation` node. + child: Rc, + /// What to compute from it. + evaluation: PopulationStatistic, + }, + // Implemented since #560 (identical child schemas). + SummaryMerge { + /// The states to merge; all have the same schema. + children: Vec>, + }, + // Reserved: migrated but unimplemented. + SummarySubtract { + left: Rc, // the state to subtract from + right: Rc, // the state to remove from `left` + }, + SummaryDelete { + summary_input: Rc, // the state + key: ColumnId, // the key column whose entries are removed + }, + SummaryJoin { + outer: Rc, // one input state + inner: Rc, // the other input state + key: ColumnId, // the join key column + family: FieldDataType, // the summary type of the result + }, + Extension { + child: Rc, // the input + name: String, // the deployment-defined operator name + }, +} + +impl ASAPOp { + /// The input nodes. For `SummaryAgg` this also includes nodes used by + /// subqueries inside its `filter`. + pub fn children(&self) -> Vec<&Rc>; + /// The same operator with each input replaced by `f(input)`. + pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Self; + /// The operator's name, e.g. `"SummaryAgg"`, for messages and display. + pub fn kind_name(&self) -> &'static str; + /// Whether the operator is reserved and cannot be built yet: Subtract, + /// Delete, Join, Extension. + pub fn is_unimplemented(&self) -> bool; + /// The summary type this operator outputs: `family` for SummaryAgg and + /// SummaryJoin, the inputs' state type for SummaryMerge, `None` otherwise. + pub fn produced_state(&self) -> Option<&FieldDataType>; + /// The output schema, derived from the operator and its inputs. An error + /// for a reserved operator or an invalid input. + pub fn output_schema(&self) -> Result; + /// The output kind: `State` for operators that output state, otherwise + /// the value kind of the input. + pub fn output_kind(&self) -> OperatorResultKind; + /// Checks the inputs (§5, "When is it rejected?"). An error if they do + /// not fit the operator. + pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>; +} +``` + +### 6.3 Schema and field types (`crates/types/src/pre_asap/schema.rs`) ```rust /// Position of a column in one schema (0-based). Local to that schema: @@ -876,7 +1007,7 @@ pub enum DataType { } ``` -### 6.2 Summary-family parameters (`crates/types/src/post_asap/sketch.rs`) +### 6.4 Summary-family parameters (`crates/types/src/post_asap/sketch.rs`) ```rust /// Which exact accumulator. None of them has parameters, so `ExactParams` @@ -971,7 +1102,7 @@ pub enum WaveletKind { Haar } pub enum WaveletParams { Haar { coeffi pub enum StatModelKind { Parametric } pub enum StatModelParams { Parametric { family: String } } ``` -### 6.3 Update input and readouts (`post_asap/sketch.rs`, `post_asap/maintained_population.rs`) +### 6.5 Update input and readouts (`post_asap/sketch.rs`, `post_asap/maintained_population.rs`) This part of the code answers two questions about summary state: @@ -1077,135 +1208,11 @@ pub struct CurrentSeriesInput { A finalized value's accuracy statement is `ResultGuarantee` (`post_asap/guarantee.rs`): `metric` (which error is measured, e.g. rank error), `bound` (the error bound), `failure_probability` (the chance the bound does not hold) and `provenance` (which estimates the bound came from). It is attached to readout and finalized nodes, never to raw state. -### 6.4 ASAP operators (`crates/types/src/ir/asap.rs`) - -```rust -pub const UNIMPLEMENTED_ASAP_OP: &str = - "this ASAP operator is reserved: schema, accuracy, timing and export are not implemented"; - -pub enum ASAPOp { - SummaryAgg { - /// The input rows. - child: Rc, - /// The summary type of the output `state` field. Never `Plain`. - family: FieldDataType, - /// What each input row adds to the state (item and weight). - input: SummaryUpdate, - /// The grouping: `Reduce(by columns)`, or `PerEntity` for one state per - /// input series without grouping. - reduction: Reduction, - /// Whether each group gets its own sketch or all groups share one. - grouping: GroupingStrategy, - /// Rows to include, applied before updating the state. `None`: all rows. - filter: Option, - }, - SummaryEstimate { - /// The node that produces the sketch state. - summary_input: Rc, - /// What to read out of it. - query: SketchStatistic, - }, - FinalizeExactAccumulator { - /// The node that produces the exact accumulator state. - child: Rc, - }, - MaintainPopulation { - /// The input rows; must match `population.input`. - child: Rc, - /// What population to keep and which reads it supports. - population: MaintainedPopulation, - }, - EvaluatePopulation { - /// The `MaintainPopulation` node. - child: Rc, - /// What to compute from it. - evaluation: PopulationStatistic, - }, - // Implemented since #560 (identical child schemas). - SummaryMerge { - /// The states to merge; all have the same schema. - children: Vec>, - }, - // Reserved: migrated but unimplemented. - SummarySubtract { - left: Rc, // the state to subtract from - right: Rc, // the state to remove from `left` - }, - SummaryDelete { - summary_input: Rc, // the state - key: ColumnId, // the key column whose entries are removed - }, - SummaryJoin { - outer: Rc, // one input state - inner: Rc, // the other input state - key: ColumnId, // the join key column - family: FieldDataType, // the summary type of the result - }, - Extension { - child: Rc, // the input - name: String, // the deployment-defined operator name - }, -} - -impl ASAPOp { - /// The input nodes. For `SummaryAgg` this also includes nodes used by - /// subqueries inside its `filter`. - pub fn children(&self) -> Vec<&Rc>; - /// The same operator with each input replaced by `f(input)`. - pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Self; - /// The operator's name, e.g. `"SummaryAgg"`, for messages and display. - pub fn kind_name(&self) -> &'static str; - /// Whether the operator is reserved and cannot be built yet: Subtract, - /// Delete, Join, Extension. - pub fn is_unimplemented(&self) -> bool; - /// The summary type this operator outputs: `family` for SummaryAgg and - /// SummaryJoin, the inputs' state type for SummaryMerge, `None` otherwise. - pub fn produced_state(&self) -> Option<&FieldDataType>; - /// The output schema, derived from the operator and its inputs. An error - /// for a reserved operator or an invalid input. - pub fn output_schema(&self) -> Result; - /// The output kind: `State` for operators that output state, otherwise - /// the value kind of the input. - pub fn output_kind(&self) -> OperatorResultKind; - /// Checks the inputs (§5, "When is it rejected?"). An error if they do - /// not fit the operator. - pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>; -} -``` - -### 6.5 Summary coverage (`crates/types/src/ir/node.rs`, `crates/types/src/ir/summary_coverage.rs`) +### 6.6 Summary coverage (`crates/types/src/ir/summary_coverage.rs`) The code for §4. ```rust -/// One node of the DAG. Immutable and shared through `Rc`. -pub struct OperatorNode { - /// What the node does: an ordinary operator or an ASAP operator. - pub operator: Operator, - /// What kind of output it has: `Relation`, `InstantVector`, - /// `RangeVector`, or `State` (summary state). Derived from `operator`. - pub result_kind: OperatorResultKind, - /// The output columns. Derived from `operator` and its children. - pub schema: Schema, - /// The accuracy statement of the output, once known. `None` does not - /// mean exact. - pub guarantee: Option, - /// When the node runs: `IngestionTime` or `QueryTime`. `None` until - /// planning assigns it. - pub timing: Option, - /// Cache for `coverage()`. Filled on first use, never serialized, - /// ignored by equality, emptied on clone. Not a source of truth: - /// coverage can always be derived again. - coverage_cache: CoverageCache, -} - -impl OperatorNode { - /// The node's coverage (§4.2), derived on first use and then cached. - /// `Some` for a `SummaryAgg` and a valid `SummaryMerge`; `None` for every - /// other node. - pub fn coverage(&self) -> Option<&SummaryCoverage>; -} - pub struct SummaryCoverage { /// What the state computes: the `SummaryAgg` and its sub-DAG, with the /// conditions that went into `selection` taken out (§4.2.2). From 502ccfbf1d0fa5c43c6db635adb2518ef3116bc7 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:14:53 +0000 Subject: [PATCH 52/59] =?UTF-8?q?docs:=20add=20the=20cost=20of=20deriving?= =?UTF-8?q?=20coverage=20(=C2=A74.2.3)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 34 +++++++++++++++++++ 1 file changed, 34 insertions(+) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 66dc8dca0..6fe9f0acb 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -114,6 +114,7 @@ A node in the physical data will represent the data or summary instance, so a no | What is it based on? | Goldstein & Larson view matching (SIGMOD 2001), explained with an example | §4.1 | | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2.1 | | How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.2.2 | +| How expensive is it? | proportional to the few operators directly under the `SummaryAgg`, not to the whole sub-DAG; cached per node | §4.2.3 | | What uses it? | merge, rollup, slice, reuse, subtract | §5.6 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.1, §6.6 | | What is left out? | the deployment and runtime implementation, e.g. SDS | §4.3 | @@ -380,6 +381,39 @@ Any other shape stays in `definition`: In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column. +#### 4.2.3 Cost of deriving coverage + +**When it runs.** `coverage()` derives a node's coverage the first time it is called and caches it on the node, so each node pays once. Building a `SummaryMerge` (`OperatorNode::new`) also derives its coverage once to reject an invalid merge. + +**Sizes used below.** + +| Symbol | Meaning | Typical size | +|---|---|---| +| `d` | operators on the walk: the `Filter`, `Project`, `TimeRange` and `TimeShift` nodes directly under the `SummaryAgg` (§4.2.2) | a few | +| `c` | filter conditions on the walk (after splitting at `AND`), including `SummaryAgg.filter` and `Scan.predicates` | a few | +| `f` | columns in the `SummaryAgg`'s input schema | tens | +| `v` | values in one value set (`IN` list) | a few | +| `n` | inputs of a `SummaryMerge` | panes per window, regions, … | +| `b` | boxes in one input's selection | 1 for a `SummaryAgg` | +| `N` | nodes in a `definition` | the sub-DAG size | + +**`SummaryAgg`: `O(c · (d·f + v²) + d)`.** + +- Each condition is checked once. Following its column up through the `Project`s costs `O(d·f)`; checking that the column's name is unique costs `O(f)`; building and intersecting a value set costs `O(v²)`, because values are compared by type, not hashed. +- Rebuilding the definition creates at most `d` new nodes. Everything below the walk is shared, not copied. +- So the cost depends only on the few operators directly under the `SummaryAgg`, **not on the size of the sub-DAG below them**. A `SummaryAgg` over a large join costs the same as one over a `Scan`. + +**`SummaryMerge`: `O(n·N + n²·b²·f·v² + (n·b)³)` in the worst case.** + +| Step | Cost | Why | +|---|---|---| +| inputs' coverage | 0 extra | each input's coverage is already cached | +| equal definitions | `O(n·N)` node comparisons (each compares an operator and a schema), often `O(n)` | structural comparison of each input's definition with the first one. Shared nodes (`Rc`) compare in `O(1)`, and node pairs already proven equal are remembered | +| no overlap | `O(n²·b²·f·v²)` | every pair of inputs, every pair of boxes, every shared column | +| union of selections | `O((n·b)³)` box comparisons, worst case | joins touching ranges and value sets until nothing more joins; each join restarts the scan | + +For the common cases this is small: `n` one-minute panes have one box each with no columns, so the merge costs `O(n·N)` for the definitions and `O(n²)` for overlap. Nested merges keep `n` small: a merge of merges compares only its direct inputs, whose coverage is cached. + ### 4.3 What coverage does not contain Coverage only says what a state means and which rows it took. The deployment and runtime implementation is not part of coverage: it belongs to the downstream deployment runtime, for example the summary data store (SDS), reading source data, and building, storing and serving summary instances. From 7f56d685680d41cc82d0c0393071afbca0835cf9 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:20:50 +0000 Subject: [PATCH 53/59] docs: fix correctness issues found in review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Column identity follows the SummaryAgg's input schema (a Project renames and drops the table); mark §6.1/§6.6 as #646 and per-family overlap as #592; panes come from window composition (Pass 2); add IRate; three selection shapes; current SummaryEstimate top-k output; time row of the definition table; drop unchecked paper section numbers. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 72 ++++++++++--------- 1 file changed, 40 insertions(+), 32 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 6fe9f0acb..87690b819 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -91,7 +91,7 @@ Based on our requirement, each field should contain the following information. | `Count` (approximate) | `Sketch`: `Cms`, `CountSketch`, `UnivMon` | `SummaryEstimate(PointCount { .. })` | | `TopK` | `Sketch`: `CmsWithHeap`, `CountSketchWithHeap` | `SummaryEstimate(TopK { k })` | | `FrequencyL2`, `FrequencyEntropy` | `Sketch`: `UnivMon` | `SummaryEstimate(FrequencyL2 \| FrequencyEntropy)` | - | `Sum`, `Count`, `Min`, `Max`, `Rate`, `Increase` (exact) | `ExactAggregate(ExactKind, …)` | `FinalizeExactAccumulator` | + | `Sum`, `Count`, `Min`, `Max`, `Rate`, `IRate`, `Increase` (exact) | `ExactAggregate(ExactKind, …)` | `FinalizeExactAccumulator` | The sketch candidates are `summary_candidates(intent)` in `crates/asap-aware-mapping/src/replacement.rs`; the readouts are `SketchStatistic` ([§6.5](#65-update-input-and-readouts-post_asapsketchrs-post_asapmaintained_populationrs)). @@ -115,7 +115,7 @@ A node in the physical data will represent the data or summary instance, so a no | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2.1 | | How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.2.2 | | How expensive is it? | proportional to the few operators directly under the `SummaryAgg`, not to the whole sub-DAG; cached per node | §4.2.3 | -| What uses it? | merge, rollup, slice, reuse, subtract | §5.6 | +| What uses it? | merge (in #646); rollup, slice, reuse (planned); subtract (reserved) | §5.6, §5.7 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.1, §6.6 | | What is left out? | the deployment and runtime implementation, e.g. SDS | §4.3 | @@ -165,13 +165,13 @@ GROUP BY job; | **Ranges** | for each column, the interval of values it may have | comparisons with a constant: `x > 5`, `x = 5`, `BETWEEN` | `day ∈ [1, 31]` | `day ∈ [5, 10]`, `region ∈ ['us', 'us']` | | **Residuals** | every other condition; the algorithm does not try to understand them | e.g. `a + b > 10`, `lower(name) LIKE 'a%'`, `x = 1 OR y = 2` | none | none | -**Step 2: does `V` contain every row `Q` needs?** (§3.1.2 of the paper) +**Step 2: does `V` contain every row `Q` needs?** - **Ranges:** each range of `Q` lies inside the same column's range in `V`. `day ∈ [5, 10]` is inside `[1, 31]` ✓. `V` has no range on `region`, so any `region` is in `V` ✓. - **Residuals:** each residual of `V` also appears in `Q`. Since residuals are not understood, the only safe case is when `Q` has the same condition. - **Equivalence classes:** each column equality of `V` also holds in `Q`. -**Step 3: can the answer be computed from `V`'s output?** (§3.3) +**Step 3: can the answer be computed from `V`'s output?** - **Compensating filter:** where `Q` is narrower than `V`, the extra condition is applied to `V`'s rows. So its columns must be in `V`'s output: `day` and `region` are ✓. - **Regrouping:** `Q`'s `GROUP BY` must be a subset of `V`'s. `{job}` ⊆ `{region, job, day}` ✓, so each group of `Q` is the sum of some groups of `V`. @@ -186,8 +186,8 @@ GROUP BY job; **Limits that matter for us:** -- It answers a query from **one** view. Combining several views (a union) is left out (§3.1). -- It supports only `SUM` and `COUNT`, whose groups can be added up again. +- It answers a query from **one** view. Combining several views (a union) is left out. +- It supports aggregates whose groups can be added up again: `SUM` and `COUNT` (and `AVG` computed from them). **Existing implementation.** The `WHERE` split is implemented for DataFusion in [`datafusion-contrib/datafusion-materialized-views`](https://github.com/datafusion-contrib/datafusion-materialized-views), `src/rewrite/normal_form.rs` (`SpjNormalForm`, `Predicate { eq_classes, ranges_by_equivalence_class, residuals }`). It rejects plans that contain an `Aggregate` or a `Join`. @@ -220,7 +220,7 @@ If coverage were one thing, for example the whole sub-DAG compared as a unit, `S Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` is the sub-DAG without its row filters, `σ` is the selection, and `G` the grouping. -[^gl]: **What we take from Goldstein & Larson, and what we add.** The view's tables, joins and residuals become the sub-DAG `C` below the `SummaryAgg`; the aggregate and its argument become the summary family and its input; `GROUP BY` becomes the `SummaryAgg` grouping `G`. These three are in `definition`. The paper's ranges become `selection` (§4.2.2). A compensating filter on the view's output becomes slicing, allowed only on a column of `G`, and regrouping to a smaller `GROUP BY` becomes rollup (§5.6). We add three things: **unions of states** (the paper uses one view at a time; `SummaryMerge` combines several, so we also check that their selections do not overlap), **summary families** (the paper only re-adds `SUM` and `COUNT`; each family says how its inputs may overlap, §5.6), and **more kinds of conditions** (value sets, hash partitions, and time relative to the evaluation time, §4.2.2). +[^gl]: **What we take from Goldstein & Larson, and what we add.** The view's tables, joins and residuals become the sub-DAG `C` below the `SummaryAgg`; the aggregate and its argument become the summary family and its input; `GROUP BY` becomes the `SummaryAgg` grouping `G`. These three are in `definition`. The paper's ranges become `selection` (§4.2.2). A compensating filter on the view's output becomes slicing, allowed only on a column of `G`, and regrouping to a smaller `GROUP BY` becomes rollup (§5.6). We add three things: **unions of states** (the paper uses one view at a time; `SummaryMerge` combines several, so we also check that their selections do not overlap), **summary families** (the paper only re-aggregates `SUM` and `COUNT`; each family says how its inputs may overlap, §5.6), and **more kinds of conditions** (value sets, hash partitions, and time relative to the evaluation time, §4.2.2). #### 4.2.2 Deriving the definition and the selection @@ -261,7 +261,8 @@ The `definition` is the `SummaryAgg` together with its sub-DAG, with every condi | a `Filter` whose conditions all move into `selection` | removed | | a `Filter` with some conditions that stay | kept, with only the conditions that stay | | `Scan.predicates` and `SummaryAgg.filter` | trimmed the same way | -| a range `TimeRange` over a `TimeShift` that becomes relative time | removed | +| the one range `TimeRange`, and any `TimeShift`s, when they become relative time | removed | +| a `TimeShift` with no range `TimeRange`, or two or more range `TimeRange`s | unchanged: no time is taken out | | any other operator | unchanged | So the `definition` holds what the state computes: the computation `C` with its remaining conditions, the summary family and its parameters, the input column, and the grouping `G`. @@ -310,11 +311,11 @@ In the worked example: | Condition | Found at | Rule 1: same rows at the `SummaryAgg`? | Rule 2: simple? | Result | |---|---|---|---|---| | `value < 100` | `SummaryAgg.filter` | yes, it is already there | yes, an interval | `selection`: `value ∈ (−∞, 100)` | -| `region = 'us'` | `Filter` | yes: `Project` passes `region` through (renamed `r`) | yes, a value set | `selection`: `m.region ∈ {us}` | +| `region = 'us'` | `Filter` | yes: `Project` passes `region` through (renamed `r`) | yes, a value set | `selection`: `r ∈ {us}` | | `value * 2 > 10` | `Filter` | yes | **no**: it is on an expression, not a column | stays in `definition` | | 1 minute, shifted by 2 | `TimeRange` + `TimeShift` | yes | yes, relative time | `selection`: `(−3m, −2m]` | -So the `selection` is `value ∈ (−∞, 100)`, `m.region ∈ {us}`, time `(−3m, −2m]`. +So the `selection` is `value ∈ (−∞, 100)`, `r ∈ {us}`, time `(−3m, −2m]`. **Rule 1 in detail.** Imagine moving the condition up, one operator at a time, until it is just below the `SummaryAgg`. Every operator it passes must leave the picked rows unchanged. Whether it can pass depends on what the operator does: @@ -331,7 +332,7 @@ A condition *above* `rate` has nothing to pass. `Filter(value > 0, rate(...))` k (This is filter pushdown in reverse. DataFusion's `PushDownFilter` uses the same rules to move filters down.) -**Rule 2 in detail.** `selection` can hold only two shapes of condition, each on a single column: a set of values, or a range. Hash partitions will be a third shape later. +**Rule 2 in detail.** `selection` can hold only three shapes of condition, each on a single column: allowed values, forbidden values, or a range. Hash partitions will be a fourth shape later. | Shape | Written as | Example | Stored as | |---|---|---|---| @@ -351,9 +352,10 @@ Any other shape stays in `definition`: **Which column a condition is on.** -- A column is named by its source table and name, `(table, name)`: in a join, `shipping.region = 'us'` and `billing.region = 'us'` are different conditions. -- A rename keeps the original name: `region AS r` is still `m.region`. -- If two output columns have the same `(table, name)` (for example `Project [a AS k, b AS k]`), a condition on `k` cannot tell them apart and stays in `definition`. +- A column is named by its table and name, `(table, name)`, **as the `SummaryAgg` reads it** (in the schema of its child). Directly above a join, `shipping.region = 'us'` and `billing.region = 'us'` are different conditions. +- A `Project` gives its columns new names and drops the table: after `region AS r`, the column is `(none, r)`, so the condition becomes `r ∈ {us}`. Since the `definition` contains the same `Project`, two states that rename the same way still compare equal. +- If two columns the `SummaryAgg` reads have the same `(table, name)` (for example `Project [a AS k, b AS k]`), a condition on `k` cannot tell them apart and stays in `definition`. +- PromQL labels have no table, so a label is named by its name alone. - Values of different types are never treated as different: `1` and `1.0` might be equal, so `x = 1` and `x = 1.0` are treated as possibly overlapping. **Time.** There are two kinds: @@ -364,7 +366,7 @@ Any other shape stays in `definition`: | **Absolute** | an interval on the timestamp column | `ts >= t0 AND ts < t1` → `ts ∈ [t0, t1)` | - PromQL windows exclude their start, so relative windows are open on the left. -- Stage 2 builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). +- Window composition (Pass 2 of logical optimization) builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). - An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. @@ -466,7 +468,7 @@ All examples read a table, so values are `Relation`. For PromQL series they woul coverage() of the SummaryAgg ┌──────────────────────────────────────────────────────────────────┐ │ definition: this SummaryAgg over Scan t (the Filter removed) │ -│ selection: t.region ∈ {us}, t.latency ∈ (−∞, 10000) │ +│ selection: region ∈ {us}, latency ∈ (−∞, 10000) │ └──────────────────────────────────────────────────────────────────┘ ``` @@ -507,7 +509,7 @@ coverage() of the SummaryAgg | Question | Answer | |---|---| -| What comes out? | the same columns, with `state` replaced by the answer: `quantile` Float64 here. Counts and cardinalities are Int64. A top-k readout instead returns the top rows themselves (the same shape as an exact `Sort` + `Limit`) | +| What comes out? | the same columns, with `state` replaced by the answer: `quantile` Float64 here. Counts and cardinalities are Int64 (Float64 when the state was built per series, `PerEntity`). A top-k readout is a `topk` Utf8 field today; #579 changes it to return the top rows themselves (the same shape as an exact `Sort` + `Limit`) | | When is it rejected? | the input is not a sketch, or the sketch cannot answer the question. For example, asking a KLL for a cardinality | | What is its coverage? | none: the output is a value | | State or value? | value. The node carries the readout's error bound | @@ -568,7 +570,7 @@ coverage() of the SummaryAgg | Question | Answer | |---|---| | What comes out? | the same columns as the input, all plain. Only the result kind `State` marks it as maintained | -| When is it rejected? | the population description does not match the child. For table rows: the child must be that same table `Scan`, the value column a non-null Float64, and the grouping valid. For PromQL series: a scan of the same metric, labels and grouping, under an instant `TimeRange` | +| When is it rejected? | the population description does not match the child. For table rows: the child must be that same table `Scan`, the value column a non-null Float64, and the grouping valid. For PromQL series: a scan of the same metric, labels and grouping, under an instant `TimeRange` of the lookback (which may be left out for the default 5-minute lookback) | | What is its coverage? | none. Maintained populations are not merged today | | State or value? | state, because it must also track membership changes. It is read only through `EvaluatePopulation` | @@ -619,14 +621,14 @@ coverage() of the SummaryAgg SummaryMerge { children: Vec, group_by: Reduction } ``` -**Example A: time panes.** PromQL `quantile_over_time(0.99, m[2m])` built from two one-minute panes, the way Stage 2 builds them. Both panes read the same `Scan`. +**Example A: time panes.** The p99 of all samples of metric `m` over the last 2 minutes (one KLL over every series), built from two one-minute panes, the way window composition (Pass 2) builds them. Both panes read the same `Scan`. ```text ( next operator ) ▲ │ State: state Sketch(KLL k=200) │ - [[ SummaryMerge ]] group_by: nothing + [[ SummaryMerge ]] ▲ selection: time (−2m, 0] │ ┌─────────────────┴─────────────────┐ @@ -664,7 +666,7 @@ merge (────────────────────── ▲ │ State: job Utf8, state Sketch(KLL k=200) │ - [[ SummaryMerge ]] group_by: by job + [[ SummaryMerge ]] ▲ selection: region ∈ {us, eu} │ ┌─────────────────┴─────────────────┐ @@ -735,18 +737,19 @@ input groups output groups | `A` + a KLL with `k = 400` | ✗ | different definitions (parameters) | | | `(A + B)` + a state over `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | -**Whether inputs may overlap** depends on the summary family (`FieldDataType::family_merges` and `merge_relation`, #592): +**Whether inputs may overlap** depends on the summary family. This is #592 (open), which adds `FieldDataType::family_merges` and `merge_relation`; #646 alone requires every merge to be overlap-free: | Rule | Families | Example | |---|---|---| | **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | -| **may overlap** | HLL, exact `Min`/`Max`, distinct sets | HLL over `A`'s and `C`'s rows ✓: a value seen twice is still one distinct value; the result covers `(−90s, 0]` | +| **may overlap** | HLL, exact `Min`/`Max` | HLL over `A`'s and `C`'s rows ✓: a value seen twice is still one distinct value. The result keeps both boxes, `(−1m, 0]` and `(−90s, −30s]`, which together cover `(−90s, 0]` | +| **cannot merge** | `CmsWithHeap`, `CountSketchWithHeap`, Theta, KMV, exact `Rate`/`IRate`/`Increase`, samples, wavelets, models | no merge is defined for them | | **right inside left** | subtraction | 5.7 | | Question | Answer | |---|---| | What comes out? | the children's schema; with `group_by`, only the remaining group columns | -| When is it rejected? | no children; a child is not state; the definitions differ (different column, parameters, filters or source); the selections overlap where the family does not allow it; `group_by` is not a subset of the children's grouping | +| When is it rejected? | no children; a child is not state; the definitions differ (different column, parameters, filters or source); the selections may overlap (#646; with #592, only where the family does not allow it); the family cannot merge at all (#592); `group_by` is not a subset of the children's grouping (planned) | | What is its coverage? | the shared `definition` (with the new grouping), and the union of the children's selections. Touching ranges join; gaps stay as separate pieces | | State or value? | state in, state out | @@ -814,6 +817,8 @@ OperatorNode §6.1 Every operator in a DAG is wrapped in an `OperatorNode`. `coverage()` is explained in §6.6. +Shown as of #646. On `main` today the node still has a declared `pub coverage: Option` field (with `with_coverage()` and `requires_coverage()`); #646 replaces it with the derived `coverage()` below. + ```rust /// One node of the DAG. Immutable and shared through `Rc`. pub struct OperatorNode { @@ -846,6 +851,8 @@ impl OperatorNode { ### 6.2 ASAP operators (`crates/types/src/ir/asap.rs`) +In the code `ASAPOp>` is generic over how it refers to its inputs (`C` can also be a node id). It is shown here with `C = Rc`. + ```rust pub const UNIMPLEMENTED_ASAP_OP: &str = "this ASAP operator is reserved: schema, accuracy, timing and export are not implemented"; @@ -918,8 +925,9 @@ impl ASAPOp { /// The input nodes. For `SummaryAgg` this also includes nodes used by /// subqueries inside its `filter`. pub fn children(&self) -> Vec<&Rc>; - /// The same operator with each input replaced by `f(input)`. - pub fn map_children(&self, f: impl FnMut(&Rc) -> Rc) -> Self; + /// The same operator with each input replaced by `f(input)`; `f` may + /// change how inputs are referred to (e.g. `Rc` to a node id). + pub fn map_children(&self, f: impl FnMut(&Rc) -> D) -> ASAPOp; /// The operator's name, e.g. `"SummaryAgg"`, for messages and display. pub fn kind_name(&self) -> &'static str; /// Whether the operator is reserved and cannot be built yet: Subtract, @@ -1054,7 +1062,7 @@ pub struct SketchKind { /// Which aggregation intents the sketch answers: `Quantile` (quantiles), /// `Cardinality` (distinct counts), `Frequency` (item counts), `TopK` /// (heavy hitters), or `Universal` (frequency moments such as L2 and - /// entropy, plus counts and distinct counts). Derived from `algorithm`. + /// entropy, plus counts, distinct counts and top-k). Derived from `algorithm`. category: SketchCategory, /// Which sketch algorithm. algorithm: SketchAlgorithm, @@ -1149,7 +1157,7 @@ This part of the code answers two questions about summary state: |---|---|---|---|---| | p99 of `latency` | KLL | none: KLL has no keys | `Column(latency)`: the value itself | `UnknownOrSigned` | | how often each `endpoint` occurs | Count-Min | `Column(endpoint)` | `Constant(1.0)`: each row counts once | `NonNegative(UnitCount)` | -| `topk(5, rate(http_requests_total[5m]))` | Count-Min with heap | `Tuple(label columns)`: one item per series | `Column(value)`: the rate | `NonNegative(ResetAwareCounterDerivative)` | +| `topk(5, rate(http_requests_total[5m]))` | Count-Min with heap | `Tuple(every column except value and the group columns)`: one item per series, including its timestamp | `Column(value)`: the rate | `NonNegative(ResetAwareCounterDerivative)` | `weight_domain` matters because some sketches (e.g. Count-Min) are only accurate when weights are never negative. The planner records why a weight is non-negative; if it cannot prove it, the weight counts as possibly negative. @@ -1240,11 +1248,11 @@ pub struct CurrentSeriesInput { } ``` -A finalized value's accuracy statement is `ResultGuarantee` (`post_asap/guarantee.rs`): `metric` (which error is measured, e.g. rank error), `bound` (the error bound), `failure_probability` (the chance the bound does not hold) and `provenance` (which estimates the bound came from). It is attached to readout and finalized nodes, never to raw state. +A finalized value's accuracy statement is `ResultGuarantee` (`post_asap/guarantee.rs`): `metric` (which error is measured, e.g. rank error), `bound` (the error bound), `failure_probability` (the chance the bound does not hold) and `provenance` (which estimates the bound came from). It is attached to readout and finalized nodes, and to `MaintainPopulation` nodes (whose population is exact); never to sketch or accumulator state. ### 6.6 Summary coverage (`crates/types/src/ir/summary_coverage.rs`) -The code for §4. +The code for §4, as of #646. On `main` today `SummaryCoverage` is still the older declared form (`source` plus `regions`). ```rust pub struct SummaryCoverage { @@ -1290,12 +1298,12 @@ pub enum Constraint { impl SummaryCoverage { /// Computes the coverage of a summary node from its sub-DAG (§4.2.2). /// Errors: `NotSummary` for a node that is not a `SummaryAgg` or - /// `SummaryMerge`; for a merge, `EmptyMerge` (no inputs), + /// `SummaryMerge`, or a merge input that has no coverage; for a merge, `EmptyMerge` (no inputs), /// `DefinitionMismatch` (inputs compute different things) or /// `PossibleOverlap` (inputs may share rows). pub fn derive(node: &OperatorNode) -> Result; } ``` -- `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections the family does not allow), but it does not store the result. +- `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections that may overlap; #592 relaxes the overlap check per family), but it does not store the result. - A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. From 63510de5576df7d98b2063c62a56c1fe398fb110 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 9 Oct 2026 21:46:52 +0000 Subject: [PATCH 54/59] docs: fix second-review findings Merge coverage is cached at construction (#646); full per-type overlap table from #592; k mismatch is a schema rejection; #579 top-k shape; tighter cost bounds; walk-order caveat for combining conditions; Project qualifier; absolute time marked later; Example B scope; one-walk wording; define by[job] and closed schema; link Pass 2; window notation. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 61 ++++++++++--------- 1 file changed, 31 insertions(+), 30 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 87690b819..5b93c6f05 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -93,7 +93,7 @@ Based on our requirement, each field should contain the following information. | `FrequencyL2`, `FrequencyEntropy` | `Sketch`: `UnivMon` | `SummaryEstimate(FrequencyL2 \| FrequencyEntropy)` | | `Sum`, `Count`, `Min`, `Max`, `Rate`, `IRate`, `Increase` (exact) | `ExactAggregate(ExactKind, …)` | `FinalizeExactAccumulator` | - The sketch candidates are `summary_candidates(intent)` in `crates/asap-aware-mapping/src/replacement.rs`; the readouts are `SketchStatistic` ([§6.5](#65-update-input-and-readouts-post_asapsketchrs-post_asapmaintained_populationrs)). + The sketch candidates are `summary_candidates(intent)` in `crates/asap-aware-mapping/src/replacement.rs`; exact accumulators come from `exact_realization` there (`Count`) and from `function_rules.rs` (the others); the readouts are `SketchStatistic` ([§6.5](#65-update-input-and-readouts-post_asapsketchrs-post_asapmaintained_populationrs)). - **Time window aggregation intents**: whether states built over smaller windows can answer a larger one. This depends on how the family combines states: - **Merge** (`SummaryMerge`, §5.6): states over disjoint panes combine into the state of their union, e.g. two 1-minute KLL states answer a 2-minute quantile. Requires a mergeable family. @@ -113,7 +113,7 @@ A node in the physical data will represent the data or summary instance, so a no | Why not put it in the schema? | States worth merging cover different data but must have the same schema | below | | What is it based on? | Goldstein & Larson view matching (SIGMOD 2001), explained with an example | §4.1 | | What does it store? | `definition` (what is computed) + `selection` (which rows were taken) | §4.2.1 | -| How is it computed? | by the planner, from the sub-DAG the node covers: first the definition, then the selection | §4.2.2 | +| How is it computed? | by the planner, from the sub-DAG the node covers, in one walk that yields both the definition and the selection | §4.2.2 | | How expensive is it? | proportional to the few operators directly under the `SummaryAgg`, not to the whole sub-DAG; cached per node | §4.2.3 | | What uses it? | merge (in #646); rollup, slice, reuse (planned); subtract (reserved) | §5.6, §5.7 | | Where is it in the code? | `OperatorNode::coverage()`, `SummaryCoverage::derive` | §6.1, §6.6 | @@ -127,7 +127,7 @@ A node in the physical data will represent the data or summary instance, so a no | | State A | State B | Equal? | |---|---|---|---| | schema | `(job: Utf8, state: KLL{k=200})` | `(job: Utf8, state: KLL{k=200})` | yes, so the merge is allowed | - | what it summarizes | time `[0,1)` | time `[1,2)` | no, which is why merging them is useful | + | what it summarizes | time `(0, 1m]` | time `(1m, 2m]` | no, which is why merging them is useful | - If coverage were part of the schema, these two schemas would differ and the merge would be rejected. The only merge left would be a state with an exact copy of itself, which counts every observation twice. @@ -202,7 +202,7 @@ A summary state is a stored aggregation, like `V` in §4.1, whose aggregate is a | **`definition`** | *What* is computed? | | **`selection`** | *Which rows* went in? | -**Example.** Three states over table `t`, all with the same schema `(job Utf8, state Sketch(KLL k=200))`: +**Example.** Three states over table `t`, all with the same schema `(job Utf8, state Sketch(KLL k=200))`. `KLL(latency) by[job]` means a KLL sketch of `latency` for each `job`: | State | Sub-DAG | `definition` | `selection` | |---|---|---|---| @@ -305,6 +305,7 @@ In the worked example, `value < 100`, `region = 'us'` and the time window move i 3. **Putting the two rules together.** - If both answers are yes: take the condition out of the sub-DAG and put it into `selection`. - If either answer is no: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. + - One more case stays in `definition`: a condition on a column that already has a lifted condition it cannot be combined with (a forbidden-value set and a range, or values of different types). Conditions are taken in walk order: the `SummaryAgg.filter`, then the `Filter`s from top to bottom, then `Scan.predicates`. In the worked example: @@ -332,7 +333,7 @@ A condition *above* `rate` has nothing to pass. `Filter(value > 0, rate(...))` k (This is filter pushdown in reverse. DataFusion's `PushDownFilter` uses the same rules to move filters down.) -**Rule 2 in detail.** `selection` can hold only three shapes of condition, each on a single column: allowed values, forbidden values, or a range. Hash partitions will be a fourth shape later. +**Rule 2 in detail.** `selection` can hold only three shapes of condition, each on a single column compared with non-NULL constants (`region = 'us'` and `'us' = region` both work): allowed values, forbidden values, or a range. Hash partitions will be a fourth shape later. | Shape | Written as | Example | Stored as | |---|---|---|---| @@ -353,7 +354,7 @@ Any other shape stays in `definition`: **Which column a condition is on.** - A column is named by its table and name, `(table, name)`, **as the `SummaryAgg` reads it** (in the schema of its child). Directly above a join, `shipping.region = 'us'` and `billing.region = 'us'` are different conditions. -- A `Project` gives its columns new names and drops the table: after `region AS r`, the column is `(none, r)`, so the condition becomes `r ∈ {us}`. Since the `definition` contains the same `Project`, two states that rename the same way still compare equal. +- A `Project` gives its columns new names and replaces their table with its own qualifier (none by default): after `region AS r`, the column is `(none, r)`, so the condition becomes `r ∈ {us}`. Since the `definition` contains the same `Project`, two states that rename the same way still compare equal. - If two columns the `SummaryAgg` reads have the same `(table, name)` (for example `Project [a AS k, b AS k]`), a condition on `k` cannot tell them apart and stays in `definition`. - PromQL labels have no table, so a label is named by its name alone. - Values of different types are never treated as different: `1` and `1.0` might be equal, so `x = 1` and `x = 1.0` are treated as possibly overlapping. @@ -363,10 +364,10 @@ Any other shape stays in `definition`: | Kind | Comes from | Example | |---|---|---| | **Relative** to the evaluation time | a range `TimeRange(w)` over a `TimeShift(s)` → `(−(s+w), −s]` | `TimeRange(1m)` alone → `(−1m, 0]`; over `TimeShift(1m)` → `(−2m, −1m]` | -| **Absolute** | an interval on the timestamp column | `ts >= t0 AND ts < t1` → `ts ∈ [t0, t1)` | +| **Absolute** (later) | an interval on the timestamp column | `ts >= t0 AND ts < t1` → `ts ∈ [t0, t1)` | - PromQL windows exclude their start, so relative windows are open on the left. -- Window composition (Pass 2 of logical optimization) builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). +- Window composition ([Pass 2](planner-layering.md#pass-2-asap-aware-common-subexpression-elimination) of logical optimization) builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). - An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. @@ -385,7 +386,7 @@ In the last two rows `TimeRange(5m)` stays in `definition`: it is the input wind #### 4.2.3 Cost of deriving coverage -**When it runs.** `coverage()` derives a node's coverage the first time it is called and caches it on the node, so each node pays once. Building a `SummaryMerge` (`OperatorNode::new`) also derives its coverage once to reject an invalid merge. +**When it runs.** `coverage()` derives a node's coverage the first time it is called and caches it on the node, so each node pays once. Building a `SummaryMerge` (`OperatorNode::new`) derives its coverage to reject an invalid merge and caches the result, so `coverage()` and `validate_structure` do not derive it again. `validate_structure` derives it only for a merge that was not built through `new`. **Sizes used below.** @@ -393,28 +394,28 @@ In the last two rows `TimeRange(5m)` stays in `definition`: it is the input wind |---|---|---| | `d` | operators on the walk: the `Filter`, `Project`, `TimeRange` and `TimeShift` nodes directly under the `SummaryAgg` (§4.2.2) | a few | | `c` | filter conditions on the walk (after splitting at `AND`), including `SummaryAgg.filter` and `Scan.predicates` | a few | -| `f` | columns in the `SummaryAgg`'s input schema | tens | +| `f` | columns in the widest schema on the walk | tens | | `v` | values in one value set (`IN` list) | a few | | `n` | inputs of a `SummaryMerge` | panes per window, regions, … | | `b` | boxes in one input's selection | 1 for a `SummaryAgg` | | `N` | nodes in a `definition` | the sub-DAG size | -**`SummaryAgg`: `O(c · (d·f + v²) + d)`.** +**`SummaryAgg`: `O(c · (d·f + v²) + d·f)`.** - Each condition is checked once. Following its column up through the `Project`s costs `O(d·f)`; checking that the column's name is unique costs `O(f)`; building and intersecting a value set costs `O(v²)`, because values are compared by type, not hashed. -- Rebuilding the definition creates at most `d` new nodes. Everything below the walk is shared, not copied. +- Rebuilding the definition creates at most `d` new nodes, each copying a schema (`O(f)`). Everything below the walk is shared, not copied. - So the cost depends only on the few operators directly under the `SummaryAgg`, **not on the size of the sub-DAG below them**. A `SummaryAgg` over a large join costs the same as one over a `Scan`. **`SummaryMerge`: `O(n·N + n²·b²·f·v² + (n·b)³)` in the worst case.** | Step | Cost | Why | |---|---|---| -| inputs' coverage | 0 extra | each input's coverage is already cached | -| equal definitions | `O(n·N)` node comparisons (each compares an operator and a schema), often `O(n)` | structural comparison of each input's definition with the first one. Shared nodes (`Rc`) compare in `O(1)`, and node pairs already proven equal are remembered | +| inputs' coverage | at most once per input | usually already cached; otherwise derived and cached now | +| equal definitions | `O(n·N·f)`: `n·N` node comparisons, each comparing an operator and a schema; often `O(n·f)` | structural comparison of each input's definition with the first one. Shared nodes (`Rc`) compare in `O(1)`, and node pairs already proven equal are remembered | | no overlap | `O(n²·b²·f·v²)` | every pair of inputs, every pair of boxes, every shared column | | union of selections | `O((n·b)³)` box comparisons, worst case | joins touching ranges and value sets until nothing more joins; each join restarts the scan | -For the common cases this is small: `n` one-minute panes have one box each with no columns, so the merge costs `O(n·N)` for the definitions and `O(n²)` for overlap. Nested merges keep `n` small: a merge of merges compares only its direct inputs, whose coverage is cached. +For the common cases this is small: `n` one-minute panes in time order have one box each with no columns, so the merge costs `O(n·N·f)` for the definitions, `O(n²)` for overlap and `O(n²)` for the union. Nested merges keep `n` small: a merge of merges compares only its direct inputs, whose coverage is cached. ### 4.3 What coverage does not contain @@ -442,7 +443,7 @@ This section walks through each summary operator with one small example. For eac | `( next operator )` | whatever consumes the result | | `selection: …` next to a node | the `selection` part of that node's `coverage()` | -All examples read a table, so values are `Relation`. For PromQL series they would be `InstantVector`. +All examples read a table, so values are `Relation`. For PromQL series they would be `InstantVector`. A *closed schema* lists every column of the table. For readability, a filter on a table is drawn as a `Filter` over the `Scan`; the frontend folds it into `Scan.predicates`, which gives the same coverage (§4.2.2). ### 5.1 `SummaryAgg`: values → state @@ -474,8 +475,8 @@ coverage() of the SummaryAgg | Question | Answer | |---|---| -| What comes out? | the group columns (`job`), then one field `state` whose type is the summary type, here `Sketch(KLL k=200)`. Each `job` appears once | -| When is it rejected? | the summary type is a plain value type; the input is already state; the input column (`latency`) is not in the child's schema; `filter` is not a boolean | +| What comes out? | the group columns (`job`), then one field `state` whose type is the summary type, here `Sketch(KLL k=200)`. Each `job` appears once. With `PerEntity` (one state per series), the input columns are kept and `state` replaces the value column | +| When is it rejected? | the summary type is a plain value type; the input is already state; the input or item column (`latency`) is not in the child's schema; `filter` is not a boolean | | What is its coverage? | always present. Both `Filter` conditions are simple, so they move into `selection`, and the `definition` is the `SummaryAgg` over the bare `Scan t`. A condition like `latency * 2 > 10` would stay in the `definition` (§4.2.2) | | State or value? | state. This is where values become state, so the result has no error bound yet | @@ -509,7 +510,7 @@ coverage() of the SummaryAgg | Question | Answer | |---|---| -| What comes out? | the same columns, with `state` replaced by the answer: `quantile` Float64 here. Counts and cardinalities are Int64 (Float64 when the state was built per series, `PerEntity`). A top-k readout is a `topk` Utf8 field today; #579 changes it to return the top rows themselves (the same shape as an exact `Sort` + `Limit`) | +| What comes out? | the same columns, with `state` replaced by the answer: `quantile` Float64 here. Counts and cardinalities are Int64 (Float64 when the state was built per series, `PerEntity`). A top-k readout is a `topk` Utf8 field today; #579 changes it to return the selected rows: the partition keys, the item identity columns and a Float64 `value` score | | When is it rejected? | the input is not a sketch, or the sketch cannot answer the question. For example, asking a KLL for a cardinality | | What is its coverage? | none: the output is a value | | State or value? | value. The node carries the readout's error bound | @@ -603,7 +604,7 @@ coverage() of the SummaryAgg | Question | Answer | |---|---| | What comes out? | the same shape as an ordinary `Aggregate` by the grouping: `quantile_0_99` Float64 here, or `sum`, `count`, `avg`. Top-k instead returns the selected rows | -| When is it rejected? | the population was not set up for the question: `quantiles` must be on for a quantile, and `k` must be at most `max_k` for top-k | +| When is it rejected? | the child is not a `MaintainPopulation`, or the population was not set up for the question: `quantiles` must be on for a quantile (and `q` finite), and `k` must be at most `max_k` for top-k | | What is its coverage? | none | | State or value? | value | @@ -659,7 +660,7 @@ pane 0 (─────────────] merge (───────────────────────────] ``` -**Example B: regions.** The US and EU states of 5.1 merge into one state for both regions: +**Example B: regions.** `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1, without the latency filter) merge into one state for both regions: ```text ( next operator ) @@ -734,23 +735,23 @@ input groups output groups | `A + C` | ✗ | `(−60s, −30s]` is in both, so those rows would be counted twice | | | `A + A` | ✗ | every row is in both | | | `A + D` | ✗ | different definitions: `D` only has rows with `value * 2 > 10` | | -| `A` + a KLL with `k = 400` | ✗ | different definitions (parameters) | | +| `A` + a KLL with `k = 400` | ✗ | different schema: `k` is part of the state type, so the merge is rejected before coverage is checked | | | `(A + B)` + a state over `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | -**Whether inputs may overlap** depends on the summary family. This is #592 (open), which adds `FieldDataType::family_merges` and `merge_relation`; #646 alone requires every merge to be overlap-free: +**Whether inputs may overlap** depends on the summary type (family and algorithm, §3). This is #592 (open), which adds `FieldDataType::family_merges` and `merge_relation`; #646 alone requires every merge to be overlap-free: -| Rule | Families | Example | +| Rule | Summary types | Example | |---|---|---| -| **must not overlap** | counting families: KLL, Count-Min, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | +| **must not overlap** | KLL, DDSketch, Count-Min, Count Sketch, UnivMon, exact `Sum`/`Count` | KLL `A + C` ✗: the rows in `(−60s, −30s]` would be counted twice | | **may overlap** | HLL, exact `Min`/`Max` | HLL over `A`'s and `C`'s rows ✓: a value seen twice is still one distinct value. The result keeps both boxes, `(−1m, 0]` and `(−90s, −30s]`, which together cover `(−90s, 0]` | -| **cannot merge** | `CmsWithHeap`, `CountSketchWithHeap`, Theta, KMV, exact `Rate`/`IRate`/`Increase`, samples, wavelets, models | no merge is defined for them | +| **cannot merge (yet)** | `CmsWithHeap`, `CountSketchWithHeap`, Theta, KMV, exact `Rate`/`IRate`/`Increase`, samples, wavelets, models | no sound merge is modeled yet, so #592 rejects them | | **right inside left** | subtraction | 5.7 | | Question | Answer | |---|---| | What comes out? | the children's schema; with `group_by`, only the remaining group columns | -| When is it rejected? | no children; a child is not state; the definitions differ (different column, parameters, filters or source); the selections may overlap (#646; with #592, only where the family does not allow it); the family cannot merge at all (#592); `group_by` is not a subset of the children's grouping (planned) | -| What is its coverage? | the shared `definition` (with the new grouping), and the union of the children's selections. Touching ranges join; gaps stay as separate pieces | +| When is it rejected? | no children; a child is not state; the schemas differ (e.g. different sketch parameters); the definitions differ (different column, filters or source); the selections may overlap (#646; with #592, only where the family does not allow it); the family cannot merge at all (#592); `group_by` is not a subset of the children's grouping (planned) | +| What is its coverage? | the shared `definition` (with the new grouping, once `group_by` exists), and the union of the children's selections. Touching ranges join; gaps stay as separate pieces | | State or value? | state in, state out | **Other uses of coverage.** The planner also uses coverage to read or reuse a state without merging: @@ -790,7 +791,7 @@ OperatorNode §6.1 │ ├── SummaryAgg │ │ ├── family: FieldDataType ───────────┐ §6.3 → §6.4 │ │ └── input: SummaryUpdate │ §6.5 (what each row adds) -│ ├── SummaryEstimate.query: SketchStatistic│ §6.5 (what is read out) +│ ├── SummaryEstimate.query: SketchStatistic │ §6.5 (what is read out) │ ├── MaintainPopulation.population │ §6.5 │ └── EvaluatePopulation.evaluation │ §6.5 ├── schema: Schema │ §6.3 @@ -1305,5 +1306,5 @@ impl SummaryCoverage { } ``` -- `OperatorNode::new` rejects an invalid `SummaryMerge` (different definitions, or selections that may overlap; #592 relaxes the overlap check per family), but it does not store the result. +- `OperatorNode::new` and `validate_structure` reject an invalid `SummaryMerge` (different definitions, or selections that may overlap; #592 relaxes the overlap check per family). The derived coverage is cached on the node. - A `SummaryAgg` always has coverage: what cannot go into `selection` stays in `definition`. From afb81c2995d5ca2b6adbfa32475a0b19c92997d9 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 10 Oct 2026 01:05:35 +0000 Subject: [PATCH 55/59] =?UTF-8?q?docs:=20fix=20=C2=A71=20grammar;=20start?= =?UTF-8?q?=20every=20example=20from=20its=20PromQL=20or=20SQL=20query?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 225 +++++++++++------- 1 file changed, 143 insertions(+), 82 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 5b93c6f05..09cc1a691 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -1,13 +1,13 @@ # Schema and Physical Data for ASAP Primitives -This document is the single source of truth for the schema, and column design for ASAP Primitives. This is used in the logical stage (LogicalASAPDAG), and physical stage (PhysicalASAPDAG). +This document is the single source of truth for the schema and column design of ASAP primitives. The design is used in the logical stage (LogicalASAPDAG) and the physical stage (PhysicalASAPDAG). ## 1. Goal, problem, and requirements -Unlike existing Database engines, which work on raw data or explicitly defined materialized tables with schema and column names provided by the users, ASAPPlanner is designed for querying and execution over the mix of raw data and ASAP Primitives. ASAP primitives are usually compact summaries over raw data. Therefore, it introduces new requirement when we design the schema and node definitions for LogicalASAPDAG and PhysicalASAPDAG. +Existing database engines work on raw data, or on materialized tables whose schema and column names the user defines explicitly. ASAPPlanner instead plans and executes queries over a mix of raw data and ASAP primitives, which are usually compact summaries of raw data. This adds new requirements to the design of the schema and the node definitions of LogicalASAPDAG and PhysicalASAPDAG. -Assuming we have the Logical DAG defined for a canonicalized representation for a batch of queries (the `LogicalDAG` produced by the frontends, [planning stages §0](planner-layering.md#0-language-specific-frontends); the stages that follow are in [planner-layering.md](planner-layering.md#stages)). -The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG ([decoupling operators from scalar expressions](decoupling_op_and_expr.md), with the unified operator type in [operator sharing §1.1](operator-sharing.md#11-unified-operator-type)), but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: +We assume a Logical DAG that represents a batch of queries in canonical form (the `LogicalDAG` produced by the frontends, [planning stages §0](planner-layering.md#0-language-specific-frontends); the stages that follow are in [planner-layering.md](planner-layering.md#stages)). +The LogicalASAPDAG shares and reuses the NonASAP operator and ScalarExpr nodes of the LogicalDAG ([decoupling operators from scalar expressions](decoupling_op_and_expr.md), with the unified operator type in [operator sharing §1.1](operator-sharing.md#11-unified-operator-type)), but replaces some of its operators with operators that work on ASAP primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: | Operator | Input → output | Status | |---|---|---| @@ -23,20 +23,18 @@ The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in | `Extension` | state → state, named by an extension | reserved | §5 walks through each of them. -Each of the Summary operators also require the ASAP primitive information above to inter-operate correctly, preserving semantic correctness. +Each summary operator also requires information about the ASAP primitives it reads, so that the operators work together and the query keeps its meaning. -Basically, the following information should be represented to preserve the equivalent query semantics when we introduce ASAP Primitives to logical query representation, and following physical one. +So when ASAP primitives are introduced into the logical query representation, and then into the physical one, the following information must be represented to keep the query semantics equivalent: -- What type of the ASAP Primitive is -- What is the ASAP Primitive parameters -- What data sources a ASAP primitive summarizes -- What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents +- What type of ASAP primitive it is +- What its parameters are +- What data sources it summarizes +- What query intents it can support, e.g. statistical aggregation intents and time window aggregation intents +This information is combined with the information of relational and time series operators, such as group by / reduction, filtering, projection, join and time series selection. - -And these information will be combined with relational or time series query operator information, such as group by/reduction, filtering, projection, join, time series selection, together. - -Therefore, these requirements drive the following schema and metadata, node information, and column design. +These requirements drive the schema, metadata, node and column design below. @@ -61,11 +59,11 @@ Two consequences for the design: - Existing systems keep aggregate state internal to one operator. ASAPPlanner makes it a first-class column type so that one state can be shared, merged and stored across queries, which is what §3 and §4 add. ## 3. Proposed schema design -Schema represents the **metadata** of information flow along an **edge** between two nodes in a logical or physical DAG. The schema field is associated with the node in the DAG. The consumer of the node in the DAG takes the schema from the producer node as input. +A schema is the **metadata** of the data that flows along an **edge** between two nodes in a logical or physical DAG. Each node stores the schema of its output, and the node that consumes it takes that schema as its input. -Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. -Based on our requirement, each field should contain the following information. -1. **What type of the ASAP Primitive is.** The field's type is a [`FieldDataType`](#63-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: +The schema definition is shared by LogicalDAG, LogicalASAPDAG and PhysicalASAPDAG. A schema contains fields, and each field maps to a column in the physical data representation. +Based on the requirements in §1, each field contains the following information. +1. **What type of ASAP primitive it is, and its parameters.** The field's type is a [`FieldDataType`](#63-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: - **Raw value**: `Plain(DataType)`, e.g., a number or a string. - **Summary state**: described from coarse to fine by four levels: @@ -80,7 +78,7 @@ Based on our requirement, each field should contain the following information. - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.4](#64-summary-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. - So a quantile KLL sketch with `k = 200`, one instance per group, has the type `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, PerSubpopulationInstance)`. -2. **What query intent the summarized ASAP Primitive can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: +2. **What query intents it can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: - **Statistical aggregation intents** (`AggIntent`, `crates/types/src/pre_asap/agg_intent.rs`): which aggregate the state can answer, and how it is read out. @@ -104,7 +102,7 @@ Based on our requirement, each field should contain the following information. ## 4. Proposed Node field design -A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. This field is the node's **coverage**. +A node represents data or a summary instance, so a node records **what data sources an ASAP primitive summarizes**. This is the node's **coverage**. **At a glance** @@ -122,7 +120,7 @@ A node in the physical data will represent the data or summary instance, so a no **Why coverage is not part of the schema.** - `SummaryMerge` requires all inputs to have the same schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). -- Two summary states worth merging always cover different data. For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2: +- Two summary states worth merging always cover different data. For example, `quantile_over_time(0.99, latency[2m])` (p99 of each series of metric `latency`, whose series are identified by label `job`) can be answered from two one-minute KLL states, one for minute 0–1 and one for minute 1–2 (§5.6, Example A): | | State A | State B | Equal? | |---|---|---|---| @@ -202,7 +200,18 @@ A summary state is a stored aggregation, like `V` in §4.1, whose aggregate is a | **`definition`** | *What* is computed? | | **`selection`** | *Which rows* went in? | -**Example.** Three states over table `t`, all with the same schema `(job Utf8, state Sketch(KLL k=200))`. `KLL(latency) by[job]` means a KLL sketch of `latency` for each `job`: +**Example.** Three queries over table `t`: + +```sql +-- S_us: p99 latency per job, US rows +SELECT job, approx_percentile_cont(latency, 0.99) FROM t WHERE region = 'us' GROUP BY job; +-- S_eu: p99 latency per job, EU rows +SELECT job, approx_percentile_cont(latency, 0.99) FROM t WHERE region = 'eu' GROUP BY job; +-- S_size: p99 size per job, EU rows +SELECT job, approx_percentile_cont(size, 0.99) FROM t WHERE region = 'eu' GROUP BY job; +``` + +With an error target, the planner answers each quantile from a KLL state. All three states have the same schema `(job Utf8, state Sketch(KLL k=200))`. `KLL(latency) by[job]` means a KLL sketch of `latency` for each `job`: | State | Sub-DAG | `definition` | `selection` | |---|---|---|---| @@ -227,28 +236,30 @@ Formally, a state means `family(input(σ(C)))` for each group of `G`, where `C` The planner computes the coverage of a node from the sub-DAG the node covers, not from a declaration. One walk down the sub-DAG produces both parts: every filter condition either moves into `selection` or stays in `definition`. The definition is described first, then how the planner decides which conditions move into the selection. -**Worked example.** A KLL of request latency per job, over metric `m` with columns `region`, `job`, `value`: +**Worked example.** p99 latency per job over table `t` (columns `job`, `region`, `latency`): + +```sql +SELECT job, + approx_percentile_cont(latency, 0.99) FILTER (WHERE latency < 100) +FROM (SELECT job, region AS r, latency + FROM t + WHERE region = 'us' AND latency * 2 > 10) +GROUP BY job; +``` + +The SQL frontend lowers it as follows: `FILTER (WHERE …)` becomes the aggregate's filter, `region AS r` becomes a `Project`, and the `WHERE` conditions are folded into the `Scan`'s predicates. With an error target, the planner answers the quantile from a KLL `SummaryAgg`, read out by a `SummaryEstimate`. The sub-DAG under that `SummaryEstimate`: ```text - ( next operator ) - ▲ - │ - [[ SummaryAgg ]] KLL(value) by job, filter: value < 100 + ( SummaryEstimate p99 ) ▲ │ - [ Project ] job, region AS r, value + [[ SummaryAgg ]] KLL(latency) by job, filter: latency < 100 ▲ │ - [ Filter ] region = 'us' AND value * 2 > 10 + [ Project ] job, region AS r, latency ▲ │ - [ TimeRange ] 1m (range) - ▲ - │ - [ TimeShift ] 2m - ▲ - │ - [ Scan m ] + [ Scan t ] predicates: region = 'us' AND latency * 2 > 10 ``` ##### The definition @@ -267,22 +278,19 @@ The `definition` is the `SummaryAgg` together with its sub-DAG, with every condi So the `definition` holds what the state computes: the computation `C` with its remaining conditions, the summary family and its parameters, the input column, and the grouping `G`. -In the worked example, `value < 100`, `region = 'us'` and the time window move into `selection` (below), and `value * 2 > 10` stays: +In the worked example, `latency < 100` and `region = 'us'` move into `selection` (below), and `latency * 2 > 10` stays: ```text - ( next operator ) - ▲ - │ - [[ SummaryAgg ]] KLL(value) by job ← filter removed + ( SummaryEstimate p99 ) ▲ │ - [ Project ] job, region AS r, value ← unchanged + [[ SummaryAgg ]] KLL(latency) by job ← filter removed ▲ │ - [ Filter ] value * 2 > 10 ← region = 'us' removed + [ Project ] job, region AS r, latency ← unchanged ▲ │ - [ Scan m ] ← TimeRange, TimeShift removed + [ Scan t ] predicates: latency * 2 > 10 ← region = 'us' removed ``` **When two definitions are equal.** Merging (§5.6) requires equal definitions. @@ -301,7 +309,7 @@ In the worked example, `value < 100`, `region = 'us'` and the time window move i 1. **Collect the conditions.** Go down the sub-DAG from the `SummaryAgg` and collect every filter condition: from `Filter` nodes, from `Scan.predicates`, and from the `SummaryAgg`'s own `filter`. A condition `A AND B` counts as two conditions, `A` and `B`. 2. **Ask two questions about each condition:** - **Rule 1: would it pick the same rows if it were moved to just below the `SummaryAgg`?** `region = 'us'` below a `Project` that only renames columns: yes. `value > 5` below `rate`: no, because it filters the raw samples that `rate` reads, which changes the rate values. - - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `value * 2 > 10` is not: it is on an expression. + - **Rule 2: is it a simple condition on one column?** That is, a value set such as `region IN ('us', 'eu')`, or a range such as `latency < 100`. `latency * 2 > 10` is not: it is on an expression. 3. **Putting the two rules together.** - If both answers are yes: take the condition out of the sub-DAG and put it into `selection`. - If either answer is no: leave it in the sub-DAG, so it is part of `definition`. The paper calls such conditions *residuals*. @@ -311,12 +319,11 @@ In the worked example: | Condition | Found at | Rule 1: same rows at the `SummaryAgg`? | Rule 2: simple? | Result | |---|---|---|---|---| -| `value < 100` | `SummaryAgg.filter` | yes, it is already there | yes, an interval | `selection`: `value ∈ (−∞, 100)` | -| `region = 'us'` | `Filter` | yes: `Project` passes `region` through (renamed `r`) | yes, a value set | `selection`: `r ∈ {us}` | -| `value * 2 > 10` | `Filter` | yes | **no**: it is on an expression, not a column | stays in `definition` | -| 1 minute, shifted by 2 | `TimeRange` + `TimeShift` | yes | yes, relative time | `selection`: `(−3m, −2m]` | +| `latency < 100` | `SummaryAgg.filter` | yes, it is already there | yes, an interval | `selection`: `latency ∈ (−∞, 100)` | +| `region = 'us'` | `Scan.predicates` | yes: `Project` passes `region` through (renamed `r`) | yes, a value set | `selection`: `r ∈ {us}` | +| `latency * 2 > 10` | `Scan.predicates` | yes | **no**: it is on an expression, not a column | stays in `definition` | -So the `selection` is `value ∈ (−∞, 100)`, `r ∈ {us}`, time `(−3m, −2m]`. +So the `selection` is `latency ∈ (−∞, 100)`, `r ∈ {us}`. A PromQL example with a time window is under **Time** below. **Rule 1 in detail.** Imagine moving the condition up, one operator at a time, until it is just below the `SummaryAgg`. Every operator it passes must leave the picked rows unchanged. Whether it can pass depends on what the operator does: @@ -329,7 +336,7 @@ So the `selection` is `value ∈ (−∞, 100)`, `r ∈ {us}`, time `(−3m, − | `rate` or a window function (later) | only if it uses series labels | a label is the same for every sample of a series | below `rate(...)`: `job = 'api'` ✓; `value > 5` ✗, dropping raw samples changes the rate | | any other operator, e.g. `Join`, `Limit` | no | | | -A condition *above* `rate` has nothing to pass. `Filter(value > 0, rate(...))` keeps the rate outputs above 0, and those are exactly the rows the `SummaryAgg` reads, so it becomes `value ∈ (0, ∞)` in `selection`. Only the `TimeRange(5m)` under `rate` stays in `definition`. +A `Filter` *above* `rate` has nothing to pass: it keeps some rate outputs, and those are exactly the rows the `SummaryAgg` reads. A PromQL comparison such as `rate(m[5m]) > 0` is not lowered to a `Filter`, though, but to a comparison operator (`BinaryOp`), which the walk does not enter. So today it stays in `definition` (last row of **More examples** below). (This is filter pushdown in reverse. DataFusion's `PushDownFilter` uses the same rules to move filters down.) @@ -346,7 +353,7 @@ Any other shape stays in `definition`: | Condition | Why it is not one of the shapes | |---|---| -| `value * 2 > 10` | it is on an expression, not a column | +| `latency * 2 > 10` | it is on an expression, not a column | | `a = b` | it compares two columns | | `region = 'us' OR job = 'api'` | it uses two columns | | `name LIKE 'web%'` | it is a pattern, not a set of values or a range | @@ -359,7 +366,7 @@ Any other shape stays in `definition`: - PromQL labels have no table, so a label is named by its name alone. - Values of different types are never treated as different: `1` and `1.0` might be equal, so `x = 1` and `x = 1.0` are treated as possibly overlapping. -**Time.** There are two kinds: +**Time.** Example: `quantile_over_time(0.99, m[1m] offset 2m)` lowers to `TimeRange(1m)` over `TimeShift(2m)` over `Scan m`. Its KLL state (one per series) has `definition` = the `SummaryAgg` over `Scan m`, and `selection` = time `(−3m, −2m]`, the minute that ended 2 minutes before evaluation. There are two kinds of time: | Kind | Comes from | Example | |---|---|---| @@ -372,17 +379,17 @@ Any other shape stays in `definition`: - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. -**More examples** of what goes where: +**More examples** of what goes where. Each query's quantile is answered from a KLL `SummaryAgg`: -| Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | -|---|---|---| -| `Filter(region = 'us', Scan t)` | `Scan t` | `region ∈ {us}` | -| `Filter(latency < 100, Scan t)` | `Scan t` | `latency ∈ (−∞, 100)` | -| `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | the last 3 to 2 minutes before evaluation, `(−3m, −2m]` | -| `Filter(job = 'api', rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` | `job ∈ {api}` | -| `Filter(value * 2 > 10, rate(TimeRange(5m, Scan m)))` | `rate(TimeRange(5m, Scan m))` and the condition `value * 2 > 10` | nothing | +| Query | Sub-DAG below `SummaryAgg` | `definition` keeps | `selection` takes | +|---|---|---|---| +| SQL `… FROM t WHERE region = 'us'` | `Scan t {region = 'us'}` | `Scan t` | `region ∈ {us}` | +| SQL `… FROM t WHERE latency < 100` | `Scan t {latency < 100}` | `Scan t` | `latency ∈ (−∞, 100)` | +| PromQL `quantile_over_time(0.99, m[1m] offset 2m)` | `TimeRange(1m, TimeShift(2m, Scan m))` | `Scan m` | the last 3 to 2 minutes before evaluation, `(−3m, −2m]` | +| PromQL `quantile(0.99, rate(m{job="api"}[5m]))` | `rate(TimeRange(5m, Scan m {job = 'api'}))` | all of it | nothing: `job = 'api'` is under `rate`, which the walk cannot pass yet (Rule 1, later) | +| PromQL `quantile(0.99, rate(m[5m]) > 0)` | `BinaryOp(>, rate(TimeRange(5m, Scan m)), 0)` | all of it | nothing: a PromQL comparison is a `BinaryOp`, not a `Filter` | -In the last two rows `TimeRange(5m)` stays in `definition`: it is the input window of `rate` and changes the rate values, so it does not just pick rows. `value * 2 > 10` stays too, because it is a condition on an expression, not on a column. +`Scan t {…}` is a `Scan` with those predicates. In the last two rows `TimeRange(5m)` stays in `definition` in any case: it is the input window of `rate` and changes the rate values, so it does not just pick rows. #### 4.2.3 Cost of deriving coverage @@ -449,7 +456,15 @@ All examples read a table, so values are `Relation`. For PromQL series they woul **Operator definition.** Turns rows into summary state: one state per group. -**Example.** p99 latency by job, from a KLL sketch with `k = 200`, over table `t`, using only US rows with latency under 10 s. +**Example.** p99 latency by job, using only US rows with latency under 10 s: + +```sql +SELECT job, approx_percentile_cont(latency, 0.99) +FROM t WHERE region = 'us' AND latency < 10000 +GROUP BY job; +``` + +With an error target, the planner answers the quantile from a KLL sketch with `k = 200`: a `SummaryAgg` builds the state, and a `SummaryEstimate` (5.2) reads p99 from it. ```text ( next operator ) @@ -484,7 +499,14 @@ coverage() of the SummaryAgg **Operator definition.** Reads a number out of a sketch, for example a quantile or a count. -**Example.** Read p99 and p50 from the state in 5.1. One state feeds both readouts. +**Example.** Two queries over the same rows as 5.1, one for p99 and one for p50: + +```sql +SELECT job, approx_percentile_cont(latency, 0.99) FROM t WHERE region = 'us' AND latency < 10000 GROUP BY job; +SELECT job, approx_percentile_cont(latency, 0.5) FROM t WHERE region = 'us' AND latency < 10000 GROUP BY job; +``` + +Both need the same KLL state, so the planner builds it once (sub-DAG sharing, [Pass 2](planner-layering.md#pass-2-asap-aware-common-subexpression-elimination)) and reads it twice. ```text ( next operator ) ( next operator ) @@ -519,7 +541,13 @@ coverage() of the SummaryAgg **Operator definition.** Turns an exact accumulator (sum, count, min, max, rate, …) into its final value. -**Example.** Total bytes by host, with an exact `Sum` accumulator. +**Example.** Total bytes by host: + +```sql +SELECT host, SUM(bytes) FROM t GROUP BY host; +``` + +The planner answers `SUM` with an exact `Sum` accumulator: a `SummaryAgg` builds it, and `FinalizeExactAccumulator` reads the total. ```text ( next operator ) @@ -554,7 +582,18 @@ coverage() of the SummaryAgg **Operator definition.** Keeps every value of a population (not a sketch), so that exact quantiles and top-k can be computed later, and tracks rows entering and leaving. -**Example.** Keep all latencies per job, so that p99 and top-10 can be computed later (5.5). +**Example.** Exact p99 latency per job, and the top-10 latencies per job: + +```sql +SELECT job, percentile_cont(latency, 0.99) FROM t GROUP BY job; -- exact target +SELECT job, latency +FROM (SELECT job, latency, + ROW_NUMBER() OVER (PARTITION BY job ORDER BY latency DESC) AS rn + FROM t) +WHERE rn <= 10; +``` + +With an exact target, the planner can keep all latencies per job in one maintained population, and answer both queries from it (5.5): ```text ( next operator ) @@ -579,7 +618,7 @@ coverage() of the SummaryAgg **Operator definition.** Computes an exact statistic from a maintained population. -**Example.** p99 and the top-10 latencies by job, both from the one population in 5.4. +**Example.** The two queries of 5.4 (exact p99 per job, top-10 latencies per job), both read from the one population: ```text ( next operator ) ( next operator ) @@ -622,21 +661,28 @@ coverage() of the SummaryAgg SummaryMerge { children: Vec, group_by: Reduction } ``` -**Example A: time panes.** The p99 of all samples of metric `m` over the last 2 minutes (one KLL over every series), built from two one-minute panes, the way window composition (Pass 2) builds them. Both panes read the same `Scan`. +**Example A: time panes.** p99 of each series of metric `m` over the last 2 minutes: + +```promql +quantile_over_time(0.99, m[2m]) +``` + +Window composition (Pass 2) can answer it from two one-minute panes, one KLL per series in each, merged. Both panes read the same `Scan`. ```text ( next operator ) ▲ - │ State: state Sketch(KLL k=200) + │ State: series labels, state Sketch(KLL k=200) │ [[ SummaryMerge ]] ▲ selection: time (−2m, 0] │ ┌─────────────────┴─────────────────┐ - │ State: state Sketch(KLL k=200) │ State: state Sketch(KLL k=200) + │ State: series labels, │ State: series labels, + │ state Sketch(KLL k=200) │ state Sketch(KLL k=200) │ │ [[ SummaryAgg ]] pane 0 [[ SummaryAgg ]] pane 1 - KLL k=200, by nothing KLL k=200, by nothing + KLL k=200, per series KLL k=200, per series selection: time (−1m, 0] selection: time (−2m, −1m] ▲ ▲ │ │ @@ -660,7 +706,15 @@ pane 0 (─────────────] merge (───────────────────────────] ``` -**Example B: regions.** `KLL(latency) by[job]` for `region = 'us'` and for `region = 'eu'` (as in 5.1, without the latency filter) merge into one state for both regions: +**Example B: regions.** Two queries for p99 latency per job, one per region, and a third for both regions: + +```sql +SELECT job, approx_percentile_cont(latency, 0.99) FROM t WHERE region = 'us' GROUP BY job; +SELECT job, approx_percentile_cont(latency, 0.99) FROM t WHERE region = 'eu' GROUP BY job; +SELECT job, approx_percentile_cont(latency, 0.99) FROM t WHERE region IN ('us', 'eu') GROUP BY job; +``` + +The first two build `KLL(latency) by[job]` states. The third can be answered by merging them into one state for both regions: ```text ( next operator ) @@ -685,7 +739,14 @@ merge (────────────────────── [ Scan t ] ``` -**Example C: rollup (planned).** A `by[region, job]` state rolled up to `by[job]`: each job's state is the merge of its per-region states. The same `by[region, job]` state also answers p99 per region and job directly. +**Example C: rollup (planned).** p99 latency per region and job, and per job: + +```sql +SELECT region, job, approx_percentile_cont(latency, 0.99) FROM t GROUP BY region, job; +SELECT job, approx_percentile_cont(latency, 0.99) FROM t GROUP BY job; +``` + +One `by[region, job]` state answers the first query directly. Rolled up to `by[job]`, it answers the second: each job's state is the merge of its per-region states. ```text ( next operator ) ( next operator ) @@ -720,23 +781,23 @@ input groups output groups (eu, web) ─┴─ merge ─────────────▶ web ``` -**Which merges are allowed.** All states below are `KLL(value) by[job] over Scan m` unless noted: +**Which merges are allowed.** Each state below is the per-series KLL that answers a PromQL query; `A`, `B` and `C` have the definition `KLL(value) per series over Scan m`: -| State | Selection | -|---|---| -| `A` | time `(−1m, 0]` | -| `B` | time `(−2m, −1m]` | -| `C` | time `(−90s, −30s]` | -| `D` | time `(−1m, 0]`, but the `definition` keeps `value * 2 > 10` | +| State | Query | Selection | +|---|---|---| +| `A` | `quantile_over_time(0.99, m[1m])` | time `(−1m, 0]` | +| `B` | `quantile_over_time(0.99, m[1m] offset 1m)` | time `(−2m, −1m]` | +| `C` | `quantile_over_time(0.99, m[1m] offset 30s)` | time `(−90s, −30s]` | +| `D` | `quantile_over_time(0.99, n[1m])` | time `(−1m, 0]`, but over metric `n`: definition `KLL(value) per series over Scan n` | | Merge | Allowed? | Why | Result's selection | |---|---|---|---| | `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (touching ranges join) | | `A + C` | ✗ | `(−60s, −30s]` is in both, so those rows would be counted twice | | | `A + A` | ✗ | every row is in both | | -| `A + D` | ✗ | different definitions: `D` only has rows with `value * 2 > 10` | | +| `A + D` | ✗ | different definitions: `D` reads metric `n`, not `m` | | | `A` + a KLL with `k = 400` | ✗ | different schema: `k` is part of the state type, so the merge is rejected before coverage is checked | | -| `(A + B)` + a state over `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | +| `(A + B)` + the state of `m[1m] offset 2m`, time `(−3m, −2m]` | ✓ | a merge has coverage like any state, so merges nest | `(−3m, 0]` | **Whether inputs may overlap** depends on the summary type (family and algorithm, §3). This is #592 (open), which adds `FieldDataType::family_merges` and `merge_relation`; #646 alone requires every merge to be overlap-free: From a0be6334f56e065e54ff7e6ed786bf070290eef2 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 10 Oct 2026 01:07:20 +0000 Subject: [PATCH 56/59] =?UTF-8?q?docs:=20restore=20the=20author's=20wordin?= =?UTF-8?q?g=20in=20=C2=A71,=20=C2=A73=20and=20=C2=A74=20openings?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 38 ++++++++++--------- 1 file changed, 20 insertions(+), 18 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 09cc1a691..dfd8a650b 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -1,13 +1,13 @@ # Schema and Physical Data for ASAP Primitives -This document is the single source of truth for the schema and column design of ASAP primitives. The design is used in the logical stage (LogicalASAPDAG) and the physical stage (PhysicalASAPDAG). +This document is the single source of truth for the schema, and column design for ASAP Primitives. This is used in the logical stage (LogicalASAPDAG), and physical stage (PhysicalASAPDAG). ## 1. Goal, problem, and requirements -Existing database engines work on raw data, or on materialized tables whose schema and column names the user defines explicitly. ASAPPlanner instead plans and executes queries over a mix of raw data and ASAP primitives, which are usually compact summaries of raw data. This adds new requirements to the design of the schema and the node definitions of LogicalASAPDAG and PhysicalASAPDAG. +Unlike existing Database engines, which work on raw data or explicitly defined materialized tables with schema and column names provided by the users, ASAPPlanner is designed for querying and execution over the mix of raw data and ASAP Primitives. ASAP primitives are usually compact summaries over raw data. Therefore, it introduces new requirement when we design the schema and node definitions for LogicalASAPDAG and PhysicalASAPDAG. -We assume a Logical DAG that represents a batch of queries in canonical form (the `LogicalDAG` produced by the frontends, [planning stages §0](planner-layering.md#0-language-specific-frontends); the stages that follow are in [planner-layering.md](planner-layering.md#stages)). -The LogicalASAPDAG shares and reuses the NonASAP operator and ScalarExpr nodes of the LogicalDAG ([decoupling operators from scalar expressions](decoupling_op_and_expr.md), with the unified operator type in [operator sharing §1.1](operator-sharing.md#11-unified-operator-type)), but replaces some of its operators with operators that work on ASAP primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: +Assuming we have the Logical DAG defined for a canonicalized representation for a batch of queries (the `LogicalDAG` produced by the frontends, [planning stages §0](planner-layering.md#0-language-specific-frontends); the stages that follow are in [planner-layering.md](planner-layering.md#stages)). +The LogicalASAPDAG will share/reuse the NonASAP operator and ScalarExpr nodes in LogicalDAG ([decoupling operators from scalar expressions](decoupling_op_and_expr.md), with the unified operator type in [operator sharing §1.1](operator-sharing.md#11-unified-operator-type)), but replacing some operators in LogicalDAG with the operators operated with ASAP Primitives. The complete list is `ASAPOp` in `crates/types/src/ir/asap.rs`: | Operator | Input → output | Status | |---|---|---| @@ -23,18 +23,20 @@ The LogicalASAPDAG shares and reuses the NonASAP operator and ScalarExpr nodes o | `Extension` | state → state, named by an extension | reserved | §5 walks through each of them. -Each summary operator also requires information about the ASAP primitives it reads, so that the operators work together and the query keeps its meaning. +Each of the Summary operators also require the ASAP primitive information above to inter-operate correctly, preserving semantic correctness. -So when ASAP primitives are introduced into the logical query representation, and then into the physical one, the following information must be represented to keep the query semantics equivalent: +Basically, the following information should be represented to preserve the equivalent query semantics when we introduce ASAP Primitives to logical query representation, and following physical one. -- What type of ASAP primitive it is -- What its parameters are -- What data sources it summarizes -- What query intents it can support, e.g. statistical aggregation intents and time window aggregation intents +- What type of the ASAP Primitive is +- What is the ASAP Primitive parameters +- What data sources a ASAP primitive summarizes +- What query intent the summarized ASAP Primitive can support, e.g., statistical aggregation intents, time window aggregation intents -This information is combined with the information of relational and time series operators, such as group by / reduction, filtering, projection, join and time series selection. -These requirements drive the schema, metadata, node and column design below. + +And these information will be combined with relational or time series query operator information, such as group by/reduction, filtering, projection, join, time series selection, together. + +Therefore, these requirements drive the following schema and metadata, node information, and column design. @@ -59,11 +61,11 @@ Two consequences for the design: - Existing systems keep aggregate state internal to one operator. ASAPPlanner makes it a first-class column type so that one state can be shared, merged and stored across queries, which is what §3 and §4 add. ## 3. Proposed schema design -A schema is the **metadata** of the data that flows along an **edge** between two nodes in a logical or physical DAG. Each node stores the schema of its output, and the node that consumes it takes that schema as its input. +Schema represents the **metadata** of information flow along an **edge** between two nodes in a logical or physical DAG. The schema field is associated with the node in the DAG. The consumer of the node in the DAG takes the schema from the producer node as input. -The schema definition is shared by LogicalDAG, LogicalASAPDAG and PhysicalASAPDAG. A schema contains fields, and each field maps to a column in the physical data representation. -Based on the requirements in §1, each field contains the following information. -1. **What type of ASAP primitive it is, and its parameters.** The field's type is a [`FieldDataType`](#63-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: +Schema definition here is shared between LogicalDAG, LogicalASAPDAG, and PhysicalASAPDAG. The schema contain fields, and each field is mapping to a column in the physical data representation. +Based on our requirement, each field should contain the following information. +1. **What type of the ASAP Primitive is.** The field's type is a [`FieldDataType`](#63-schema-and-field-types-cratestypessrcpre_asapschemars) (`crates/types/src/pre_asap/schema.rs`). A column is either a raw value or a summary state: - **Raw value**: `Plain(DataType)`, e.g., a number or a string. - **Summary state**: described from coarse to fine by four levels: @@ -78,7 +80,7 @@ Based on the requirements in §1, each field contains the following information. - For a sketch, category, algorithm and parameters are bundled as one `SketchKind` ([§6.4](#64-summary-family-parameters-cratestypessrcpost_asapsketchrs)). A sketch also records its `GroupingStrategy`: one instance per group, or one shared structure (Hydra) for all groups. - So a quantile KLL sketch with `k = 200`, one instance per group, has the type `Sketch(SketchKind { Quantile, Kll, Kll { k: 200 } }, PerSubpopulationInstance)`. -2. **What query intents it can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: +2. **What query intent the summarized ASAP Primitive can support.** This is not stored in the field: it follows from the type in item 1. There are two kinds of intent: - **Statistical aggregation intents** (`AggIntent`, `crates/types/src/pre_asap/agg_intent.rs`): which aggregate the state can answer, and how it is read out. @@ -102,7 +104,7 @@ Based on the requirements in §1, each field contains the following information. ## 4. Proposed Node field design -A node represents data or a summary instance, so a node records **what data sources an ASAP primitive summarizes**. This is the node's **coverage**. +A node in the physical data will represent the data or summary instance, so a node has a field for **What data sources a ASAP primitive summarizes**. This field is the node's **coverage**. **At a glance** From 2f5bde100df5d24e4fe28b6b12d804e3c18d4b3d Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 10 Oct 2026 01:12:20 +0000 Subject: [PATCH 57/59] docs: show ASAPOp generic over its child reference, as on main Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- .../proposals/asap-primitive-schema.md | 49 ++++++++++--------- 1 file changed, 26 insertions(+), 23 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index dfd8a650b..12c80b9ac 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -915,16 +915,16 @@ impl OperatorNode { ### 6.2 ASAP operators (`crates/types/src/ir/asap.rs`) -In the code `ASAPOp>` is generic over how it refers to its inputs (`C` can also be a node id). It is shown here with `C = Rc`. +`ASAPOp` is generic over how it refers to its inputs: `C` is `Rc` by default, and can also be a node id (e.g. in a flat DAG). ```rust pub const UNIMPLEMENTED_ASAP_OP: &str = "this ASAP operator is reserved: schema, accuracy, timing and export are not implemented"; -pub enum ASAPOp { +pub enum ASAPOp> { SummaryAgg { /// The input rows. - child: Rc, + child: C, /// The summary type of the output `state` field. Never `Plain`. family: FieldDataType, /// What each input row adds to the state (item and weight). @@ -935,65 +935,68 @@ pub enum ASAPOp { /// Whether each group gets its own sketch or all groups share one. grouping: GroupingStrategy, /// Rows to include, applied before updating the state. `None`: all rows. - filter: Option, + filter: Option>, }, SummaryEstimate { /// The node that produces the sketch state. - summary_input: Rc, + summary_input: C, /// What to read out of it. query: SketchStatistic, }, FinalizeExactAccumulator { /// The node that produces the exact accumulator state. - child: Rc, + child: C, }, MaintainPopulation { /// The input rows; must match `population.input`. - child: Rc, + child: C, /// What population to keep and which reads it supports. population: MaintainedPopulation, }, EvaluatePopulation { /// The `MaintainPopulation` node. - child: Rc, + child: C, /// What to compute from it. evaluation: PopulationStatistic, }, // Implemented since #560 (identical child schemas). SummaryMerge { /// The states to merge; all have the same schema. - children: Vec>, + children: Vec, }, // Reserved: migrated but unimplemented. SummarySubtract { - left: Rc, // the state to subtract from - right: Rc, // the state to remove from `left` + left: C, // the state to subtract from + right: C, // the state to remove from `left` }, SummaryDelete { - summary_input: Rc, // the state - key: ColumnId, // the key column whose entries are removed + summary_input: C, // the state + key: ColumnId, // the key column whose entries are removed }, SummaryJoin { - outer: Rc, // one input state - inner: Rc, // the other input state - key: ColumnId, // the join key column - family: FieldDataType, // the summary type of the result + outer: C, // one input state + inner: C, // the other input state + key: ColumnId, // the join key column + family: FieldDataType, // the summary type of the result }, Extension { - child: Rc, // the input - name: String, // the deployment-defined operator name + child: C, // the input + name: String, // the deployment-defined operator name }, } -impl ASAPOp { +impl ASAPOp { /// The input nodes. For `SummaryAgg` this also includes nodes used by /// subqueries inside its `filter`. - pub fn children(&self) -> Vec<&Rc>; + pub fn children(&self) -> Vec<&C>; /// The same operator with each input replaced by `f(input)`; `f` may - /// change how inputs are referred to (e.g. `Rc` to a node id). - pub fn map_children(&self, f: impl FnMut(&Rc) -> D) -> ASAPOp; + /// change the reference type, e.g. from `Rc` to a node id. + pub fn map_children(&self, f: impl FnMut(&C) -> D) -> ASAPOp; /// The operator's name, e.g. `"SummaryAgg"`, for messages and display. pub fn kind_name(&self) -> &'static str; +} + +impl ASAPOp { /// Whether the operator is reserved and cannot be built yet: Subtract, /// Delete, Join, Extension. pub fn is_unimplemented(&self) -> bool; From 9abddffc0f002771e723c3ec8cca676e81714f70 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 10 Oct 2026 01:16:07 +0000 Subject: [PATCH 58/59] docs: describe the union and instant-selector rules as #646 implements them Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index 12c80b9ac..a7e08cf9b 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -336,7 +336,7 @@ So the `selection` is `latency ∈ (−∞, 100)`, `r ∈ {us}`. A PromQL exampl | `Project` | only if its column is passed through unchanged (a rename is fine) | the column must still be there, with the same values, above the `Project` | `Project [job, region AS r]`: `region = 'us'` ✓, it becomes `r = 'us'`. `Project [job, value * 2 AS v2]`: `value > 5` ✗, `value` is gone | | `Aggregate` (later) | only if it uses group columns | a group column has one value per group, so filtering before or after grouping keeps the same groups | below `SUM(value) by job`: `job = 'api'` ✓; `value > 5` ✗, it changes the sums | | `rate` or a window function (later) | only if it uses series labels | a label is the same for every sample of a series | below `rate(...)`: `job = 'api'` ✓; `value > 5` ✗, dropping raw samples changes the rate | -| any other operator, e.g. `Join`, `Limit` | no | | | +| any other operator, e.g. `Join`, `Limit`, an instant `TimeRange`, a `TimeShift` with `@` | no | | | A `Filter` *above* `rate` has nothing to pass: it keeps some rate outputs, and those are exactly the rows the `SummaryAgg` reads. A PromQL comparison such as `rate(m[5m]) > 0` is not lowered to a `Filter`, though, but to a comparison operator (`BinaryOp`), which the walk does not enter. So today it stays in `definition` (last row of **More examples** below). @@ -377,7 +377,7 @@ Any other shape stays in `definition`: - PromQL windows exclude their start, so relative windows are open on the left. - Window composition ([Pass 2](planner-layering.md#pass-2-asap-aware-common-subexpression-elimination) of logical optimization) builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). -- An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. +- An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. The walk also stops there, so the label matchers of an instant selector stay in `definition` too, for now. - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. @@ -422,7 +422,7 @@ Any other shape stays in `definition`: | inputs' coverage | at most once per input | usually already cached; otherwise derived and cached now | | equal definitions | `O(n·N·f)`: `n·N` node comparisons, each comparing an operator and a schema; often `O(n·f)` | structural comparison of each input's definition with the first one. Shared nodes (`Rc`) compare in `O(1)`, and node pairs already proven equal are remembered | | no overlap | `O(n²·b²·f·v²)` | every pair of inputs, every pair of boxes, every shared column | -| union of selections | `O((n·b)³)` box comparisons, worst case | joins touching ranges and value sets until nothing more joins; each join restarts the scan | +| union of selections | `O((n·b)³)` box comparisons, worst case | joins boxes whose time windows touch, or that differ only in one column's value set, until nothing more joins; each join restarts the scan | For the common cases this is small: `n` one-minute panes in time order have one box each with no columns, so the merge costs `O(n·N·f)` for the definitions, `O(n²)` for overlap and `O(n²)` for the union. Nested merges keep `n` small: a merge of merges compares only its direct inputs, whose coverage is cached. @@ -794,7 +794,7 @@ input groups output groups | Merge | Allowed? | Why | Result's selection | |---|---|---|---| -| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (touching ranges join) | +| `A + B` | ✓ | same definition, no overlap | `(−2m, 0]` (touching time windows join) | | `A + C` | ✗ | `(−60s, −30s]` is in both, so those rows would be counted twice | | | `A + A` | ✗ | every row is in both | | | `A + D` | ✗ | different definitions: `D` reads metric `n`, not `m` | | @@ -814,7 +814,7 @@ input groups output groups |---|---| | What comes out? | the children's schema; with `group_by`, only the remaining group columns | | When is it rejected? | no children; a child is not state; the schemas differ (e.g. different sketch parameters); the definitions differ (different column, filters or source); the selections may overlap (#646; with #592, only where the family does not allow it); the family cannot merge at all (#592); `group_by` is not a subset of the children's grouping (planned) | -| What is its coverage? | the shared `definition` (with the new grouping, once `group_by` exists), and the union of the children's selections. Touching ranges join; gaps stay as separate pieces | +| What is its coverage? | the shared `definition` (with the new grouping, once `group_by` exists), and the union of the children's selections. Boxes whose time windows touch join, and so do boxes that differ only in one column's value set (`{us}` and `{eu}` give `{us, eu}`). Everything else stays as separate boxes, including gaps and touching value ranges such as `latency < 100` and `latency >= 100` | | State or value? | state in, state out | **Other uses of coverage.** The planner also uses coverage to read or reuse a state without merging: From f34fc017c6ada900c84b5ecfa375883bb708709b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Sat, 10 Oct 2026 02:00:49 +0000 Subject: [PATCH 59/59] docs: touching value ranges join; instant selectors join the later lifting rules Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW --- docs/design_docs/proposals/asap-primitive-schema.md | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/design_docs/proposals/asap-primitive-schema.md b/docs/design_docs/proposals/asap-primitive-schema.md index a7e08cf9b..d8d307e1a 100644 --- a/docs/design_docs/proposals/asap-primitive-schema.md +++ b/docs/design_docs/proposals/asap-primitive-schema.md @@ -335,8 +335,8 @@ So the `selection` is `latency ∈ (−∞, 100)`, `r ∈ {us}`. A PromQL exampl | `TimeRange` (range) or `TimeShift` (without `@`) | yes | they choose a time window, but do not change any row's values | `region = 'us'` below `TimeRange(1m)` ✓ | | `Project` | only if its column is passed through unchanged (a rename is fine) | the column must still be there, with the same values, above the `Project` | `Project [job, region AS r]`: `region = 'us'` ✓, it becomes `r = 'us'`. `Project [job, value * 2 AS v2]`: `value > 5` ✗, `value` is gone | | `Aggregate` (later) | only if it uses group columns | a group column has one value per group, so filtering before or after grouping keeps the same groups | below `SUM(value) by job`: `job = 'api'` ✓; `value > 5` ✗, it changes the sums | -| `rate` or a window function (later) | only if it uses series labels | a label is the same for every sample of a series | below `rate(...)`: `job = 'api'` ✓; `value > 5` ✗, dropping raw samples changes the rate | -| any other operator, e.g. `Join`, `Limit`, an instant `TimeRange`, a `TimeShift` with `@` | no | | | +| `rate`, a window function, or an instant `TimeRange` (later) | only if it uses series labels | a label is the same for every sample of a series | below `rate(...)`: `job = 'api'` ✓; `value > 5` ✗, dropping raw samples changes the rate. Below an instant selector: `job = 'api'` ✓; `value > 5` ✗, the latest sample with value > 5 is not the latest sample | +| any other operator, e.g. `Join`, `Limit`, a `TimeShift` with `@` | no | | | A `Filter` *above* `rate` has nothing to pass: it keeps some rate outputs, and those are exactly the rows the `SummaryAgg` reads. A PromQL comparison such as `rate(m[5m]) > 0` is not lowered to a `Filter`, though, but to a comparison operator (`BinaryOp`), which the walk does not enter. So today it stays in `definition` (last row of **More examples** below). @@ -377,7 +377,7 @@ Any other shape stays in `definition`: - PromQL windows exclude their start, so relative windows are open on the left. - Window composition ([Pass 2](planner-layering.md#pass-2-asap-aware-common-subexpression-elimination) of logical optimization) builds its tumbling panes this way: a 3-minute window as panes `(−1m, 0]`, `(−2m, −1m]`, `(−3m, −2m]`, so pane times are derived, not declared (#601). -- An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. The walk also stops there, so the label matchers of an instant selector stay in `definition` too, for now. +- An instant `TimeRange` (latest sample per series) does not pick rows by time, so it stays in `definition`. The walk also stops there, so for now the label matchers of an instant selector stay in `definition` too; lifting them is planned together with `rate` (Rule 1). - The IR cannot yet write a timestamp constant, so absolute SQL time filters stay in `definition` for now. - Absolute and relative time are never compared: a state over `(−1m, 0]` and one over `ts ∈ [t0, t1)` are treated as possibly overlapping. @@ -422,7 +422,7 @@ Any other shape stays in `definition`: | inputs' coverage | at most once per input | usually already cached; otherwise derived and cached now | | equal definitions | `O(n·N·f)`: `n·N` node comparisons, each comparing an operator and a schema; often `O(n·f)` | structural comparison of each input's definition with the first one. Shared nodes (`Rc`) compare in `O(1)`, and node pairs already proven equal are remembered | | no overlap | `O(n²·b²·f·v²)` | every pair of inputs, every pair of boxes, every shared column | -| union of selections | `O((n·b)³)` box comparisons, worst case | joins boxes whose time windows touch, or that differ only in one column's value set, until nothing more joins; each join restarts the scan | +| union of selections | `O((n·b)³)` box comparisons, worst case | joins boxes that differ in only one dimension whose union is again one constraint (touching time windows, touching value ranges, or value sets of one column), until nothing more joins; each join restarts the scan | For the common cases this is small: `n` one-minute panes in time order have one box each with no columns, so the merge costs `O(n·N·f)` for the definitions, `O(n²)` for overlap and `O(n²)` for the union. Nested merges keep `n` small: a merge of merges compares only its direct inputs, whose coverage is cached. @@ -814,7 +814,7 @@ input groups output groups |---|---| | What comes out? | the children's schema; with `group_by`, only the remaining group columns | | When is it rejected? | no children; a child is not state; the schemas differ (e.g. different sketch parameters); the definitions differ (different column, filters or source); the selections may overlap (#646; with #592, only where the family does not allow it); the family cannot merge at all (#592); `group_by` is not a subset of the children's grouping (planned) | -| What is its coverage? | the shared `definition` (with the new grouping, once `group_by` exists), and the union of the children's selections. Boxes whose time windows touch join, and so do boxes that differ only in one column's value set (`{us}` and `{eu}` give `{us, eu}`). Everything else stays as separate boxes, including gaps and touching value ranges such as `latency < 100` and `latency >= 100` | +| What is its coverage? | the shared `definition` (with the new grouping, once `group_by` exists), and the union of the children's selections. Two boxes join when they differ in only one dimension and their union is again one constraint, as in the paper's one range per column: touching time windows (`(−2m, −1m]` and `(−1m, 0]` give `(−2m, 0]`), touching value ranges (`latency < 100` and `latency >= 100` give `latency ∈ (−∞, ∞)`, which still excludes NULL), or value sets of one column (`{us}` and `{eu}` give `{us, eu}`). Gaps stay as separate boxes | | State or value? | state in, state out | **Other uses of coverage.** The planner also uses coverage to read or reuse a state without merging: