Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
# ASAPPlanner

ASAPPlanner turns SQL, PromQL, and MetricsQL query workloads into legal candidate plans that may use Approximate Streaming Analytics Primitives (ASAPs), such as sketches and exact summaries. It normalizes language-specific queries into a shared representation, then enumerates and ranks semantically equivalent alternatives. Downstream systems choose, deploy, and execute a physical plan.
ASAPPlanner turns SQL, PromQL, and MetricsQL query workloads into plans that may use Approximate Streaming Analytics Primitives (ASAPs), such as sketches and exact summaries. It normalizes language-specific queries into a shared representation, enumerates semantically equivalent alternatives, and selects the cheapest one that meets each query's accuracy target. Downstream systems deploy and execute the physical plan.

## Start here

- New to the repository? Read the [planner pipeline](docs/design_docs/concepts/planner-pipeline.md), then the [glossary](docs/design_docs/concepts/glossary.md).
- Want to run a query? Follow [Run and inspect a query](docs/user_guide_docs/run-a-query.md).
- Embedding Planner? Use the [library API guide](docs/develop_docs/library-api.md).
- Extending Planner? Start with the [ASAP-aware mapping architecture](docs/develop_docs/asap-aware-mapping-architecture.md), then [extend ASAP-aware mapping](docs/develop_docs/extend-asap-aware-mapping.md).
- Extending Planner? Start with the [ASAP-aware mapping architecture](docs/develop_docs/asap-aware-mapping-architecture.md).
- Evaluating a design? Browse the [design documentation](docs/design_docs/README.md), [developer documentation](docs/develop_docs/README.md), and [user guides](docs/user_guide_docs/README.md).

The [documentation map](docs/README.md) gives each audience a complete reading path.
Expand Down
5 changes: 2 additions & 3 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,13 +19,12 @@ Start with [ASAPPlanner input, output, and workflows](design_docs/architecture/i
for the integration boundary, nested inputs, and choice of planning workflow.

Use [Public library functions and examples](develop_docs/library-api.md) for
frontend lowering, workload search, ranking, selection and DAG assembly.
frontend lowering, planning a workload with the stage pipeline, and export.

## Extend the planner

1. Read the [ASAP-aware mapping architecture](develop_docs/asap-aware-mapping-architecture.md).
2. Consult [mapping contracts](develop_docs/asap-aware-mapping-contracts.md).
3. Follow [Extend ASAP-aware mapping](develop_docs/extend-asap-aware-mapping.md).
2. Read [local logical candidates](develop_docs/local-logical-candidates.md) for Stage 1.

## Understand a design

Expand Down
76 changes: 30 additions & 46 deletions docs/design_docs/architecture/README.md
Original file line number Diff line number Diff line change
@@ -1,50 +1,43 @@
# ASAPPlanner design overview

ASAPPlanner is a reusable planning library. It converts queries and workload
requirements into a deployment-independent space of logical Post-ASAP candidates.
It does not commit, deploy, or execute a physical plan; downstream systems such
as ASAPQuery-backend bind the candidates to physical alternatives, make the
requirements into a selected logical Post-ASAP plan, through the #509 stage
pipeline. It does not deploy or execute a physical plan; downstream systems
such as ASAPQuery-backend bind the plan to physical operators, make the
deployment-level decision, and run the selected contract.

For the integration workflow, start with [ASAPPlanner input, output, and
workflows](input-output-workflow.md). It defines inputs, `CandidateLogicalASAPDAGs`, selection
and assembly workflows, and future replanning support.
For the integration workflow, start with the
[library API](../../develop_docs/library-api.md);
[ASAPPlanner input, output, and workflows](input-output-workflow.md) defines
the workload inputs.

## Planner component flow

```mermaid
flowchart TD
W["PlanningWorkload: query demand + optional data facts"]
F["Frontend dependencies: SQL catalog or PromQL time"]
E["Strategy, accuracy model, and applicable evidence"]
PRE["Frontend lowering → canonical Pre-ASAP OperatorNode roots"]
SEARCH["Whole-workload candidate search: sharing, legality, accuracy"]
SPACE["CandidateLogicalASAPDAGs: compact logical candidate DAG space"]
RANK["Optional cost_sorted: ranked inspection view"]
SELECT["Optional global_selection + assemble_selected_dag"]
DAG["Selected logical Post-ASAP DAG"]
BACKEND["Downstream: bind physical alternatives, decide deployment, compile and execute"]
M["PlanningModels: accuracy model, calibration, deployment capabilities"]
PRE["Frontend lowering → canonical Pre-ASAP roots"]
S1["Stage 1: Pass 1 alternatives per target; Pass 2 sharing variants"]
S2["Stage 2: physical candidates (materialization)"]
S3["Stage 3: accuracy and capability checks, pricing, selection"]
PLAN["Selected logical Post-ASAP plan + selection report"]
BACKEND["Downstream: bind physical operators, deploy, compile and execute"]
W --> PRE
F --> PRE
PRE --> SEARCH
E --> SEARCH
SEARCH --> SPACE
SPACE --> RANK --> BACKEND
SPACE --> SELECT --> DAG --> BACKEND
PRE --> S1 --> S2 --> S3
M --> S3
S3 --> PLAN --> BACKEND
```

`CandidateLogicalASAPDAGs` is the output of logical candidate search. Each target's candidate set holds
alternatives and rejection reasons, but no materialization decision.
Choose between two branches: inspect candidates (optionally ranked), or select
and assemble logical DAGs. Stage 2 materialization (#509) will decide per
sub-DAG whether to materialize and whether at ingestion or query time; until
then every summary runs at query time. No branch by itself deploys or executes
a physical plan.
Known-invalid evidence rejects a logical candidate. Missing accuracy evidence
leaves a constructible candidate visible in `CandidateLogicalASAPDAGs` but uncertified; default
selection does not commit it without the required guarantee. Cost evidence can
rank eligible candidates, but it cannot establish a missing guarantee or turn
an unsupported physical alternative into a deployable plan.
Stage 1 lists alternatives without pricing them. Stage 3 rejects a candidate
whose summary estimate misses its query's accuracy target (or has no accuracy
model), that needs a capability the deployment lacks, or that exceeds its
memory budget, and selects the cheapest remaining one. Rejection reasons are
reported with the selection. Cost evidence can rank valid candidates, but it
cannot establish a missing guarantee or turn an unsupported physical
alternative into a deployable plan.

## Module map

Expand All @@ -53,7 +46,7 @@ an unsupported physical alternative into a deployable plan.
| Shared IR | `asap-types` | The unified operator IR (`ir`: one `OperatorNode` before and after ASAP optimization), schemas, workloads, guarantees, and exported plan data |
| Front-end common | `frontend-common` | Name-based `UnresolvedOp` tree shared by the front ends, and `resolve_root` into the operator IR |
| Query frontends | `frontend-sql`, `frontend-promql`, `frontend-metricsql` | Parse source languages and produce canonical Pre-ASAP queries |
| ASAP-aware mapping | `asap-logical-optimizer`, `asap-physical-optimizer`, `asap-plan-selection` | #509 Stages 1–3: candidate generation, CSE, legality and accuracy propagation; physical candidates; costing and selection |
| ASAP-aware mapping | `asap-logical-optimizer`, `asap-physical-optimizer`, `asap-plan-selection` | #509 Stages 1–3: logical alternatives and sharing; physical candidates; accuracy checks, costing and selection |
| Planner facade | `asap-planner` | Lowering dispatch, the optimization pass (`OptimizationPass`, `StagePipeline`) and `optimize` |
| Developer inspection | `devtools` | Expose planner DAGs, alternatives, decisions, and explanations for inspection |
| End-to-end validation | `integration-tests` | Verify behavior across frontends, mapping, and output IR |
Expand All @@ -66,27 +59,18 @@ requirements, the planning horizon, available materialized state, downstream
capabilities, and complete cost evidence. Missing or stale evidence must remain
explicit rather than being treated as zero.

The primary output is `CandidateLogicalASAPDAGs`; `cost_sorted` derives an optional ranked
view with index-aligned costs. Downstream may inspect compatible choices
across targets rather than assuming the first candidate is a feasible
physical workload plan. Candidates carry logical summary algorithms,
parameters, and guarantees, but no materialization decision. Rejection reasons
are retained in the candidate space.
The output is the selected plan: one logical Post-ASAP root per query
(`PlanOutput`), with the selection report — the priced candidates, the rejected
ones with their reasons, and whether the selection is guaranteed optimal.
`plan_stages` also returns Stage 1's alternatives for inspection.

ASAPQuery-backend and other downstream applications translate the candidates
ASAPQuery-backend and other downstream applications translate the selected plan
into physical alternatives. They own concrete implementations, storage layout,
placement, sharding, deployment-level cost and compatibility, final commitment,
serving, and operational feedback. Their physical planning can reorder
candidates because it has evidence that the reusable Planner does not, but it
must not silently change Planner-owned semantics.

`candidate_selection::global_selection` optionally coordinates structural choices across
targets; `GlobalSelection::assemble_selected_dag` constructs a selected semantic DAG.
Those APIs do not establish physical feasibility or a
materialization decision. See the [library guide](../../develop_docs/library-api.md#optional-whole-plan-selection-and-dag-assembly)
for the distinction. Downstream may consume candidates directly and retains
responsibility for physical commitment.

## Further reading

- [Parsing and canonicalization](parse-and-canonicalize.md)
Expand Down
6 changes: 6 additions & 0 deletions docs/design_docs/architecture/asap-aware-mapping.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# ASAP-Aware Mapping

> **Historical:** this document describes the legacy replacement search
> (`ReplacementStrategy`, `CandidateLogicalASAPDAGs`, `global_selection`), which
> was removed (#630, #635). The #509 stage pipeline replaced it; see
> [planner layering](../proposals/planner-layering.md) and the
> [library API](../../develop_docs/library-api.md).

## Overview

ASAP-aware mapping decides **whether and how a query intent can be answered using summaries instead of scanning raw data**.
Expand Down
8 changes: 7 additions & 1 deletion docs/design_docs/architecture/asap-aware-plan-search.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Candidate Plan Search

> **Historical:** this document describes the legacy replacement search
> (`ReplacementStrategy`, `CandidateLogicalASAPDAGs`, `global_selection`), which
> was removed (#630, #635). The #509 stage pipeline replaced it; see
> [planner layering](../proposals/planner-layering.md) and the
> [library API](../../develop_docs/library-api.md).

ASAP-aware mapping should consider all alternatives holistically rather than optimize prematurely.

Suppose a plan contains several independent-looking decision points:
Expand Down Expand Up @@ -91,6 +97,6 @@ order.

The [code architecture](../../develop_docs/asap-aware-mapping-architecture.md)
describes current discovery and registry behavior; the
[library guide](../../develop_docs/library-api.md#optional-whole-plan-selection-and-dag-assembly)
[library guide](../../develop_docs/library-api.md)
shows selection and its evidence boundaries. Broader optimization dimensions are
tracked in the [proposal](../proposals/asap-aware-mapping/optimizations.md).
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Evidence-dependent candidates

> **Historical:** this document describes the legacy replacement search
> (`ReplacementStrategy`, `CandidateLogicalASAPDAGs`, `global_selection`), which
> was removed (#630, #635). The #509 stage pipeline replaced it; see
> [planner layering](../proposals/planner-layering.md) and the
> [library API](../../develop_docs/library-api.md).

Audience: ASAPPlanner library integrators, especially ASAPQuery-backend.

`CandidateLogicalASAPDAGs` is a space of constructible logical alternatives, not a list of
Expand Down
25 changes: 18 additions & 7 deletions docs/design_docs/architecture/input-output-workflow.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,17 @@

## Overview

> **Status:** the candidate-space output and the ranking and selection
> workflows this document describes (`CandidateLogicalASAPDAGs`, `cost_sorted`,
> `global_selection`, `CostModel`) were removed with the legacy search (#630,
> #635). ASAPPlanner now runs the #509 stage pipeline and returns the selected
> plan: `asap_planner::e2e_plan` returns a `PlanOutput`, and
> `asap_plan_selection::plan_stages` a `StagePipelineRun` with Stage 1's
> alternatives and the selection report. See the
> [library API](../../develop_docs/library-api.md) for the current workflow.
> The input sections below still describe the current `PlanningWorkload`; the
> output and workflow sections are kept as a historical record.

This document is for library integrators such as ASAPQuery-backend, not users
submitting queries through a backend.

Expand All @@ -17,7 +28,7 @@ Post-ASAP alternatives for the workload.
| `PlanningWorkload.query_workload` | Query language and one-time/repeating query workloads | Yes |
| `PlanningWorkload.data_workload` | Data arrival and optional evidence about ingestion, cardinality, and distribution | No implicit default. Set `None` when unavailable for non-PromQL workloads; PromQL requires `Some(DataWorkload)` with a nonzero ingestion interval. |
| Frontend-specific dependencies (outside `PlanningWorkload`) | `SqlCatalog` for SQL; `now_ms` and, when needed, `HistogramCatalog` for PromQL | `SqlCatalog` is required for SQL lowering; `now_ms` is required for PromQL lowering |
| Planning models | Candidate cost/ranking and accuracy composition/checking | Used by the relevant APIs; built-in `DefaultCostModel` and `DefaultAccuracyModel` are available |
| Planning models | Accuracy checking, cost calibration and deployment capabilities (`PlanningModels`) | `PlanningModels::builtin()` uses `DefaultAccuracyModel`, an illustrative calibration and unrestricted capabilities |
| External evidence and capabilities | Domain facts, measured costs, workload statistics, and runtime support | Supply when available and when the chosen optimization depends on them; absence is not proof |

Frontend lowering and candidate search are stages within this workflow, not
Expand Down Expand Up @@ -277,8 +288,8 @@ latter cannot be fabricated by one.

| Input | Where it enters / default | Why it matters |
|---|---|---|
| Accuracy model | Target-aware search takes an `AccuracyModel`; `DefaultAccuracyModel` is available. Default strategies also use it for candidate construction. | Composes candidate guarantees and checks them against requested accuracy. The model does not itself provide missing data-domain facts. |
| Cost model | Candidate strategies and `cost_sorted`/`global_selection` use a `CostModel`; `DefaultCostModel` is available. | Ranks or selects candidates. The built-in model is not a measured deployment cost for every physical implementation. |
| Accuracy model | `PlanningModels.accuracy`; `DefaultAccuracyModel` by default. Stage 3 checks each summary estimate with it. | Derives each estimate's guarantee and checks it against the requested accuracy. The model does not itself provide missing data-domain facts. |
| Cost calibration | `PlanningModels.calibration`; `Stage3Calibration::ILLUSTRATIVE` by default. | Stage 3 prices candidates analytically; the built-in calibration is not a measured deployment cost. |
| Accuracy/domain evidence | `AccuracyEvidenceProvider`; default strategies use `NoAccuracyEvidence` when no provider is supplied. | Input ranges, nonempty populations, Top-K intervals, and similar facts can certify or rule out particular approximations. Missing facts remain unknown. |
| Measured cost evidence | Supplied through a deployment-specific cost model or physical-evidence provider when cost-based physical comparison is needed. | CPU, memory, and I/O estimates must be comparable before claiming a summary beats raw recomputation. |
| Runtime capabilities | Checked by deployment-specific providers. | Prevents choosing a maintenance/window operation the intended executor cannot implement. |
Expand All @@ -297,7 +308,7 @@ the applicable strategy, accuracy target, and helper; missing evidence is not
a blanket reason to discard unrelated candidates. For the direct DDSketch
ratio above, search retains a candidate without a proven root guarantee when
domain evidence is missing; automatic `global_selection` does not choose it.
See the [candidate-search reference](../../develop_docs/library-api.md#generate-and-rank-candidates)
See the [candidate-search reference](../../develop_docs/library-api.md)
for this backend-selection path.

---
Expand Down Expand Up @@ -414,16 +425,16 @@ the result for one query root.
|:---:|
| **Input:** [CandidateLogicalASAPDAGs](asap-aware-plan-search.md) + cost model |
| ↓ |
| **Select:** [global_selection](../../develop_docs/library-api.md#what-does-global-selection-mean) chooses compatible alternatives |
| **Select:** [global_selection](../../develop_docs/library-api.md) chooses compatible alternatives |
| ↓ |
| **Assemble:** [assemble_selected_dag(root)](../../develop_docs/library-api.md#api-definition-and-example) connects those choices for each query root |
| **Assemble:** [assemble_selected_dag(root)](../../develop_docs/library-api.md) connects those choices for each query root |
| ↓ |
| **Output:** one selected logical [Post-ASAP DAG](../concepts/post-asap-ir.md) per query root |

Each output DAG specifies the chosen operators, parameters, and accuracy
guarantees. Its root is an `Rc<OperatorNode>` (the same IR as the input,
with some nodes now ASAP operators) and carries no execution timing yet; the
[API reference](../../develop_docs/library-api.md#api-definition-and-example)
[API reference](../../develop_docs/library-api.md)
describes the function signatures and return handling.

This path selects how to compute the query, not whether summary state is
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,8 @@ What that buys:
* A new optimization algorithm can be freely implemented as a trait implementation, rather than a
rule disguised to fit a two-phase pipeline it does not share.

Unchanged: `CandidateLogicalASAPDAGs`, `cost_sorted`, `global_selection`, and the interface
[input, output, and workflows](input-output-workflow.md) describes.
The legacy API (`CandidateLogicalASAPDAGs`, `cost_sorted`, `global_selection`)
was later removed (#630, #635); the "before" example below uses it.

```text
PlanningWorkload ──lowering──▶ ParsedWorkload ──OptimizationPass──▶ PlanOutput
Expand Down Expand Up @@ -123,8 +123,8 @@ and the target beneath it. When it fails, it builds every combination if
there are at most 64, and otherwise flags `Selection::method` as not
guaranteed optimal.

`ReplacementStrategy` remains a concept of the legacy candidate search, which
the default pass no longer uses.
`ReplacementStrategy` was a concept of the legacy candidate search, which has
been removed (#635).

### 3.2 Plugging in another pass

Expand Down
13 changes: 5 additions & 8 deletions docs/design_docs/concepts/accuracy-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,14 +130,11 @@ another implementation with the same algorithm name. In particular, a named
empirical calibration is different from an arbitrary benchmark's maximum
observed error; both its confidence and applicability must remain explicit.

The current interfaces still expose general parameter proposal through
`CostModel::size_params`. Default sizing is dispatched to the estimator modules through the existing
public candidate-construction entry point. Accuracy validation is independent of those proposals. The new
source-contract path centralizes HLL sizing and guarantee derivation in
Planner's accuracy module, overriding the generic proposal when the applicable
contract is supplied. It does not yet move every algorithm's sizing interface
out of CostModel. The design boundary is that parameter proposals never grant
accuracy authority to the cost model.
Stage 1 sizes every sketch through the estimator modules
(`asap_logical_optimizer::pass1::realization::default_size_params`); no cost
model proposes parameters. Accuracy validation (Stage 3's `AccuracyModel`) is
independent of sizing. The design boundary is that parameter proposals never
grant accuracy authority to the cost model.

## Composing guarantees through a DAG

Expand Down
Loading