Skip to content

feat(ir): define compatible logical summary merges - #560

Open
zzylol wants to merge 13 commits into
mainfrom
stack/528-02b-merge-structure
Open

zzylol wants to merge 13 commits into
mainfrom
stack/528-02b-merge-structure

Conversation

@zzylol

@zzylol zzylol commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Why

Builds on #645 (generic IR, CSE) and #567 (summary coverage), both merged. Extracts structural summary merge support from #555.

Before this PR: constructing ASAPOp::SummaryMerge fails because the operation is reserved.

After this PR: two structurally compatible KLL states can form a typed logical merge with no execution-phase assignment. Merge inputs must be nonempty, carry state with exactly one state field, and have identical family/parameter/grouping schemas. Empty, raw and mismatched states fail.

This PR is structural only and does not touch SummaryCoverage. Equal schemas cannot tell a KLL over latency from a KLL over size, and cannot show that the inputs cover disjoint rows. Both are decided by summary coverage, derived for every summary node by one SummaryCoverage::derive in #646, following the design in #573.

Schema/result-kind/state identity and regression tests are included. Runtime execution, materialization and derived timing belong to the later physical scopes. Subtract/delete remain reserved.

Key code interface

SummaryMerge is a logical ASAP operator whose inputs are existing operator nodes producing partial state:

pub enum ASAPOp {
    // Other variants omitted.
    SummaryMerge {
        children: Vec<Rc<OperatorNode>>,
    },
}

It uses the same node construction and validation APIs as other operators:

let merged = OperatorNode::new_shared(Operator::ASAP(
    ASAPOp::SummaryMerge {
        children: vec![pane_a, pane_b],
    },
))?;
merged.validate_structure()?;

Here pane_a and pane_b are Rc<OperatorNode> state producers. new_shared returns Result<Rc<OperatorNode>, SchemaDerivationError> and derives the output schema/result kind. validate_structure checks the complete reachable DAG, including the producers. No timing assignment is required.

The key ASAPOp methods are:

pub fn validate_inputs(&self) -> Result<(), SchemaDerivationError>;
pub fn output_schema(&self) -> Result<Schema, SchemaDerivationError>;
pub fn output_kind(&self) -> OperatorResultKind;
pub fn produced_state(&self) -> Option<&FieldDataType>;

For SummaryMerge:

  • validate_inputs enforces the compatibility rules below.
  • output_schema validates inputs, then clones the first input's complete schema.
  • output_kind is OperatorResultKind::State.
  • produced_state returns the non-plain field type from the first input. This is a metadata accessor, not a runtime merge or an independent validation step.

Grouped merge keeps groups apart: job='api' and job='worker' stay separate states. A later SummaryEstimate is the readout boundary that turns state into a plain value.

Requirements for merging two summaries

This PR enforces these structural requirements:

Requirement Check
At least one input An empty children list is rejected. The API is n-ary; a single compatible state is structurally allowed.
Every input produces state Each child has result_kind == State; raw relations/vectors are rejected.
Exactly one state field The first schema contains exactly one non-Plain field. Identical schemas enforce this on every other input too. Plain grouping fields may accompany it.
Same committed state type Complete schema equality requires the same family, algorithm/kind, parameters and any grouping layout carried by FieldDataType. KLL k=200 and k=300, or KLL and CMS, cannot merge here.
Same grouping/output layout Field positions, plain grouping-field types, names, qualifiers and nullability must match.
Same schema metadata time_index, unique_keys and closed must match as well. Matching only the state algorithm is insufficient.
Structurally valid producers Whole-DAG validate_structure checks each producer's own contracts.

Compatibility is deliberately strict: even differently named but otherwise equivalent schemas need an explicit normalization before this interface accepts them.

These checks establish typed structural compatibility only. They do not prove that the inputs summarize the same computation or cover disjoint rows (#646), that they cover the intended population/window, that a runtime implements the family's merge operation, or that the merged result meets an accuracy requirement. Those belong to #646, the logical composition rule and subsequent physical planning/selection. For example, overlapping frequency panes must not silently double-count observations; matching schemas alone cannot establish that.

When SummaryMerge can be used

During logical planning: a composition rule can construct SummaryMerge when it needs to combine compatible partial summary states—for example, several tumbling-window KLL panes answering one larger query window, or compatible partition summaries feeding a coarser computation. The planner must establish the intended input coverage and grouping semantics. The node can be constructed and structurally validated before choosing materialization.

During execution: a physical implementation can merge the state contents once its inputs are available and the selected family/runtime supports the operation. Timing follows the materialization choice in #509. For example, a query can merge previously ingested/stored panes, or merge states rebuilt for that query. #560 itself adds no runtime kernel, storage, retention, window generation or execution-phase policy, and does not imply runtime support for every declared state family.

Source: asap.rs, with regressions in crates/types/tests/summary_merge_structure.rs.

Validation: workspace tests, cargo fmt, and workspace/all-target Clippy with warnings denied all pass.


Base: main · Next: #646 · Tracker: #528

Order: #645 (merged) → #567 (merged) → #560 → #646 → #539 → #540 → #541 → #542 → #543.

🤖 Generated with Claude Code

@zzylol
zzylol changed the base branch from main to feat/summary-coverage-contract October 3, 2026 16:07
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch 6 times, most recently from 3ab6be7 to 77628c8 Compare October 3, 2026 17:29
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch from 77628c8 to 02a6a1b Compare October 3, 2026 17:40
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch 2 times, most recently from 60a9e4f to cb00197 Compare October 3, 2026 19:45
@zzylol
zzylol marked this pull request as draft October 3, 2026 20:02
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch from cb00197 to 700838b Compare October 3, 2026 20:48
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch 2 times, most recently from f2b7b3f to 3f5d369 Compare October 6, 2026 20:02
@zzylol
zzylol marked this pull request as ready for review October 6, 2026 20:02
@zzylol
zzylol force-pushed the feat/summary-coverage-contract branch from e5e4f67 to fc8eec9 Compare October 6, 2026 20:17
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch from 3f5d369 to 52b4003 Compare October 6, 2026 20:17
zzylol added a commit that referenced this pull request Oct 6, 2026
…ts fields

Name the metadata coverage to match #560 and the docs; drop the single-variant
multiplicity and deployment-specific revision; rename grouping to reduction to
match SummaryAgg; report failures through SchemaDerivationError::Coverage; revert
the unrelated PaneCoverageError rename.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the feat/summary-coverage-contract branch 2 times, most recently from bb3bf69 to 4750e71 Compare October 6, 2026 20:34
zzylol added a commit that referenced this pull request Oct 6, 2026
* feat(ir): define joint summary observation coverage

* refactor(ir): name represented observations ObservationExtent

* refactor(ir): rename observation extent to SummaryCoverage and trim its fields

Name the metadata coverage to match #560 and the docs; drop the single-variant
multiplicity and deployment-specific revision; rename grouping to reduction to
match SummaryAgg; report failures through SchemaDerivationError::Coverage; revert
the unrelated PaneCoverageError rename.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(ir): describe coverage source as any observation data source

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ir): require coverage on summary nodes; allow regions without time bounds

validate_structure rejects a SummaryAgg without coverage (CoverageError::Missing).
CoverageRegion time bounds become optional so tabular sources without a time
column can declare coverage. Population stays trusted; #570 tracks checking it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: explain the summary coverage problem with examples

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ir): keep only time and population in SummaryCoverage

Given required coverage on summary nodes, input and reduction duplicated the
producing SummaryAgg fields; drop them along with ProducerMismatch. Type source
as Source, matching Scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(cost): rename SourceCoverage to ScanSelection

SourceCoverage names the rows a physical scan reads for cost comparison, not
which observations a summary state holds; rename it so it is not confused with
SummaryCoverage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: turn summary coverage into a design document

Move it to docs/design_docs/proposals with problem and motivation,
requirements, design, alternatives and key code interfaces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: design doc for ASAP primitive schema and summary semantics

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Update asap-primitive-schema.md

* Update asap-primitive-schema.md

* Update asap-primitive-schema.md

* Update asap-primitive-schema.md

* docs: add code interfaces and per-operator examples to the ASAP primitive schema design

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: move the ASAP primitive schema design to #573

The design document and the docs it consolidates are reviewed separately on
main. This PR keeps code, tests and the ScanSelection rename in docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(ir): say coverage time is absolute and population is Utf8 equality

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ir): carry summary coverage through CSE and flat DAGs

Coverage is part of a summary state's identity: two states over different
observations are never shared, so CSE hashes and compares it, and a flat
node keeps it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch from 52b4003 to 8f58489 Compare October 6, 2026 20:45
@zzylol
zzylol changed the base branch from feat/summary-coverage-contract to main October 6, 2026 20:45
@zzylol
zzylol requested a review from Selvomega October 6, 2026 20:47
@Selvomega

Copy link
Copy Markdown
Collaborator

Revised logical foundation 2/6 · Base: #567 · Next: #537 · Tracker: #528

Revised review order: #567 → #560 → #537 → #539 → #540 → #561.

🤖 Generated with Claude Code

This order seems to be stale

zzylol added a commit that referenced this pull request Oct 6, 2026
…low multi-source summaries

Review of #560/#646 found four problems:

1. Population names came from each node's schema, which `with_schema` may
   rename. A scan whose `tier` column is named "region" made `tier = 'eu'`
   read as `{region: eu}`, so a merge with a real `{region: us}` state was
   accepted and double-counted. Columns are now named from the Scan
   operator's own schema, and a path that renames a field leaves the
   population unknown.
2. For the same reason a merge could mix states of different columns
   (`Named("latency")` reading `size` on a renamed scan). An unknown
   population only merges with the same input, so this is rejected too.
3. `OperatorNode::map_children` dropped a SummaryAgg's coverage, so
   rebuilding a merge (e.g. in canonicalize) failed. A rebuild now keeps the
   declared time bounds and reads source and population again.
4. A SummaryAgg over two sources (a join, an IN subquery over another
   table) could never validate. It now carries no coverage and cannot be
   merged.

Adds summary_coverage_derivation.rs; the four regression tests fail
before this change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread crates/types/src/ir/asap.rs Outdated
Comment on lines +222 to +233
pub fn merged_coverage(&self) -> Result<SummaryCoverage, SchemaDerivationError> {
let ASAPOp::SummaryMerge { children } = self else {
return Err(SchemaDerivationError::InvalidScalarSignature(
"coverage merge requires SummaryMerge".into(),
));
};
let inputs = children
.iter()
.map(|child| child.coverage.clone().ok_or(CoverageError::UnknownInput))
.collect::<Result<Vec<_>, _>>()?;
Ok(SummaryCoverage::merge_disjoint(&inputs)?)
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks weird. Why merged_coverage, something specific to SummaryMerge node only, will be a method for ASAPOp, a much broader type?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, it only makes sense for SummaryMerge and had to return an error for every other operator. I replaced it with SummaryCoverage::of_merge(inputs) in summary_coverage.rs, which takes the merge's inputs directly; the three callers already have them from ASAPOp::SummaryMerge { children }.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No I still feel this unsolved. Please check my comments in code later

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 42bc2f7 and e201d5a. merged_coverage and its replacement of_merge are both gone. ASAPOp has no coverage method, and this PR derives no coverage at all: the merged node carries none. SummaryMerge validation only reads each input's existing coverage, requiring it to be present (UnknownInput) and to have the same columns (ColumnMismatch, see below). Deriving coverage for every summary node, merges included, is one SummaryCoverage::derive in #646.

Comment thread crates/types/src/ir/node.rs Outdated
Comment on lines +171 to +178
match self.asap()? {
ASAPOp::SummaryAgg {
input, reduction, ..
} => Some((input, reduction)),
ASAPOp::SummaryMerge { children } => children.first()?.summary_update(),
_ => None,
}
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this function do?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It returns what a summary state was built from: the update expression (input) and grouping (reduction) of the SummaryAgg that produced it. For a SummaryMerge it returns its first input's, which merge validation requires every input to share; for any other node it returns None.

SummaryMerge uses it to require that all inputs summarize the same expression with the same grouping. Equal schemas alone don't show that: KLL over latency and KLL over size, both grouped by job, have the same schema. I expanded the doc comment to say this.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Give me a few minutes to further understand this.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Figured out with Claude code. I feel the naming is off. I thought this is a function updating the summary

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This function is now removed (e201d5a). What it returned (the producing SummaryAgg's input and reduction) is part of SummaryCoverage instead, as its columns, next to the existing rows:

pub struct SummaryCoverage {
    pub source: Source,                 // rows
    pub regions: Vec<CoverageRegion>,   // rows: time × population
    pub input: SummaryUpdate,           // columns: = SummaryAgg.input (e.g. latency)
    pub group_by: Reduction,            // columns: = SummaryAgg.reduction (e.g. by job)
}

with_coverage rejects a SummaryAgg declaration whose columns differ from the node's own (ColumnMismatch), and a merge requires identical columns across inputs, so a KLL over latency and one over size (same schema) cannot merge.

@zzylol

zzylol commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Yes, that was stale. The description now reads: #645 (merged) → #567 (merged) → #560 → #646 → #539 → #540 → #541 → #542 → #543.

zzylol added a commit that referenced this pull request Oct 6, 2026
…low multi-source summaries

Review of #560/#646 found four problems:

1. Population names came from each node's schema, which `with_schema` may
   rename. A scan whose `tier` column is named "region" made `tier = 'eu'`
   read as `{region: eu}`, so a merge with a real `{region: us}` state was
   accepted and double-counted. Columns are now named from the Scan
   operator's own schema, and a path that renames a field leaves the
   population unknown.
2. For the same reason a merge could mix states of different columns
   (`Named("latency")` reading `size` on a renamed scan). An unknown
   population only merges with the same input, so this is rejected too.
3. `OperatorNode::map_children` dropped a SummaryAgg's coverage, so
   rebuilding a merge (e.g. in canonicalize) failed. A rebuild now keeps the
   declared time bounds and reads source and population again.
4. A SummaryAgg over two sources (a join, an IN subquery over another
   table) could never validate. It now carries no coverage and cannot be
   merged.

Adds summary_coverage_derivation.rs; the four regression tests fail
before this change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread crates/types/src/ir/summary_coverage.rs Outdated

/// Coverage of a `SummaryMerge` over `inputs`: the disjoint union of their
/// coverage. Fails when an input has none or two inputs may overlap.
pub fn of_merge(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is still bad design. Since in #646 we introduced for_summary function to calculate the coverage, and that function is effectively only working for SummaryAgg nodes.

Taking an union of #646 and this PR on coverage deriving, we can see that we have

  1. of_merge calculating coverage for SummaryMerge
  2. for_summary calculating coverage for SummaryAgg
    This is very bad for the following 2 reasons
  3. Ideally we should only have 1 function deriving coverage for all summary-generating nodes, where we have relatively consistent and unified behavior, now we have 2
  4. Even for the current 2 in these PRs, neither their name nor their behavior match

I do hate the shitty patches agents are relentlessly making.

I propose here, we intendedly not implementing the coverage deriving functionality in this PR, since this is only to introduce the SummaryMerge node, and SummaryCoverage, by now, doesn't have any automatically inferred version. Then we do that in a unified way in #646, with one func

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the review. This is a great idea.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 42bc2f7: this PR no longer derives or checks coverage. of_merge, UnknownInput and MergeOutputMismatch are removed, summary_coverage.rs matches main, and SummaryMerge only checks structure. #646 will derive coverage for every summary node through a single SummaryCoverage::derive(node, time_range).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up in e201d5a: SummaryCoverage now also records the columns a state summarizes (input, group_by), so all merge compatibility beyond the schema goes through coverage. This PR still derives nothing; callers declare coverage, and with_coverage only checks the declared columns against the SummaryAgg's own fields. #646 derives rows and columns for SummaryAgg and SummaryMerge through one SummaryCoverage::derive. The design doc in #573 is updated to match.

Comment thread crates/types/src/ir/node.rs Outdated
/// requires every input to share. `None` for any other node. Merging
/// compares it so that all inputs summarize the same expression with the
/// same grouping.
pub fn summary_update(&self) -> Option<(&SummaryUpdate, &Reduction)> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggest renaming if this function is not used to actually update the summary.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Renamed to summary_input_data in 4ff8b98. It reads the producing SummaryAgg's update expression (input, e.g. the latency column) and grouping (reduction, e.g. by job); it updates nothing. The doc comment now says this and that which rows were included is coverage, not this.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: summary_input_data is removed in e201d5a; the same information is now SummaryCoverage.input and SummaryCoverage.group_by (see the thread above).

zzylol added a commit that referenced this pull request Oct 7, 2026
SummaryCoverage records which columns a state summarizes (input,
group_by) as well as which rows. Update the SummaryAgg and SummaryMerge
examples to #560: merges compare coverage columns, carry no coverage
until #646, and merged_coverage is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 12 commits October 7, 2026 14:49
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…coverage examples

SummaryCoverage no longer repeats input/reduction, so SummaryMerge compares
them through OperatorNode::summary_update. summary_coverage_examples.rs builds
each example in docs/develop_docs/summary-coverage.md as a SummaryAgg ->
SummaryMerge plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… doc

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A merged top-k heap can miss an item that is heavy in only one input, so
CmsWithHeap, CountSketchWithHeap and UnivMon states do not merge. Also
correct the merge_disjoint doc: SummaryMerge checks schema, update,
reduction and heap families, not accuracy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`ASAPOp::merged_coverage` only applied to SummaryMerge and returned an error
for every other operator. Replace it with `SummaryCoverage::of_merge`, which
takes the merge's inputs; the callers already have them. Also explain what
`OperatorNode::summary_update` returns and why merging compares it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sketches

SummaryMerge no longer derives or checks coverage: of_merge,
UnknownInput and MergeOutputMismatch are removed and summary_coverage.rs
matches main. Coverage for all summary nodes will be derived by one
function in #646. The heap-based sketch restriction is also dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It reads the producing SummaryAgg's update expression and reduction; it
does not update the summary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SummaryCoverage now records which columns a state summarizes as well as
which rows: `input` (the SummaryAgg update expression) and `group_by`
(its reduction). with_coverage rejects a SummaryAgg declaration whose
columns differ from the node's own (ColumnMismatch), and SummaryMerge
requires coverage on every input (UnknownInput) with identical columns.
merge_disjoint checks columns too. summary_input_data is removed. A
nested SummaryMerge carries no coverage until #646, so it is rejected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/528-02b-merge-structure branch from a5eb728 to 7ce0660 Compare October 7, 2026 14:55
Remove the coverage columns (input, group_by), check_columns and the
merge's coverage checks. SummaryMerge now only checks structure: at least
one input, every input is State with one state column and an identical
schema. Whether a structurally valid merge is semantically valid (same
computation, disjoint selections) is decided by summary coverage in #646,
following the design in #573.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W7qG9aFyPij5uWsyAJCxDW
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants