Skip to content

Tracking: planner/query-engine inconsistencies under streaming_engine=precompute #531

Description

@milindsrivastava1997

Tracking issue for cases where asap-planner-rs and asap-query-engine disagree — the planner plans/accepts a query the engine can't actually execute, or punts one the engine executes anyway. Scoped to streaming_engine=precompute; PromQL and SQL only (Elastic excluded, expected to mirror SQL).

Activity

  1. milindsrivastava1997 commented on Jul 14, 2026

    @milindsrivastava1997
    ContributorAuthor

    Planner/Engine Inconsistencies (streaming_engine=precompute)

    Status: Investigation complete, fixes not yet implemented
    Touches: asap-planner-rs, asap-query-engine
    Scope: streaming_engine=precompute only (arroyo excluded). PromQL and SQL only (Elastic excluded — Elastic is expected to mimic SQL's behavior, see issue #485).


    Background

    asap-planner-rs decides what to compute (StreamingConfig.aggregation_configs) and which queries it can serve (InferenceConfig.query_configs, plus a punted_queries list for queries it explicitly refuses to plan). asap-query-engine executes queries two ways at runtime:

    1. Config-path: find_query_config_{sql,promql} looks up a query_config the planner generated for this exact query text.
    2. Capability-matching fallback: when no query_config exists, find_compatible_aggregation searches StreamingConfig.aggregation_configs directly by capability (aggregation type, window size, grouping labels, spatial filter) — no planner authorization required.

    An "inconsistency" here means: the planner's decision (plan/punt) and what the engine actually does (via either path) disagree. Both directions counted:

    • Direction A: planner punts/rejects a query, but the engine (usually via capability-matching) serves it anyway.
    • Direction B: planner plans/accepts a query, but the engine's config-path can't actually execute it (error, None, or undefined result).

    All findings below were verified with a runnable repro (temporary test code, removed afterward) — not just static code reading.


    PromQL findings

    1. Capability-matching is structurally dead for Rate/Increase/Min/Max (Direction A, root cause)

    Constructs: rate(), increase(), min_over_time()/max_over_time(), spatial min(...)/max(...) — any statistic that always gets Exact treatment (see is_approximate() in asap-common/dependencies/rs/promql_utilities/src/query_logics/enums.rs:144-152,259-267, which only lists Quantile/Sum/Count/Avg/Topk as approximate).

    • Planner: asap-planner-rs/src/planner/agg_config.rs:104-124 (build_agg_configs_for_statistics) only emits a paired DeltaSetAggregator key-aggregation when the value type is CountMinSketch/HydraKLL. Combined with promql_utilities/src/query_logics/logics.rs:25-48 (map_statistic_to_precompute_operator), Exact-treated Min/Max/Rate/Increase always map to MultipleMinMax/MultipleIncrease — types that never get a paired key aggregation.
    • Engine: asap_types/src/capability_matching.rs — is_multi_population_value_type() (enums.rs:333-343) classifies MultipleSum/MultipleMinMax/MultipleIncrease as requiring a paired key aggregation, and find_compatible_aggregation returns None whenever that pairing is absent.
    • Result: capability-matching can never resolve rate/increase/min_over_time/max_over_time/spatial min/max — the fallback path is dead code for this entire class of queries, even for a perfect structural match.
    • Repro: planned rate(reqs[2m]) and max_over_time(reqs[2m]) via Controller::generate(), built exact-matching QueryRequirements from the resulting configs, called find_compatible_aggregation directly → None in both cases.

    2. Punting is advisory, not enforced — a punted query can be silently served (Direction A)

    Construct: any statistic using DatasketchesKLL/HydraKLL (e.g. quantile_over_time), where finding #1's pairing gap doesn't apply.

    • Planner: asap-planner-rs/src/planner/promql.rs:182-217 (should_be_performant) punts low-sample-count queries (e.g. t_repeat_ms / data_ingestion_interval_ms < 60); asap-planner-rs/src/promql/generator.rs:72-91 adds them to punted_queries and skips creating a query_config/aggregation_config for them. Nothing downstream enforces the punt: asap-query-engine's query_tracker/tracker.rs only logs punted_queries.len().
    • Engine: asap-query-engine/src/engines/simple_engine/promql.rs:1025-1056 — when find_query_config misses, capability-matching kicks in, matching purely on metric/statistic/window-divisibility/grouping-labels/spatial-filter, with zero awareness of punting or of the requesting query's own frequency.
    • Result: if a punted query shares metric/statistic/grouping/spatial-filter with an accepted sibling whose window size evenly divides it, the engine serves the punted query anyway — defeating the reason it was punted (protecting against low-fidelity computation).
    • Repro: planned two quantile_over_time(0.9, reqs[Xm]) variants, one punted (10 samples), one accepted (60 samples, window 60000ms). Confirmed punted_queries contains the first and it has no query_config. Called build_query_execution_context_promql on the punted query text directly against the real engine → Some(...) (served).

    Checked and discarded (PromQL)

    • Hypothesis: avg_over_time's Count leg hits a capability-matching type mismatch — did not repro. Avg/Sum/Count-over-time are always Approximate-treated and correctly map to CountMinSketch, which is in compatible_agg_types(Count) and does get a paired key aggregation.
    • histogram_quantile is unsupported by both planner and engine (no pattern exists on either side) — consistent, not a bug.

    SQL findings

    3. Capability-matching rejects every grouped COUNT/SUM the planner actually builds (Direction B, most severe — shared root cause with #4)

    Construct: any COUNT(...)/SUM(...) ... GROUP BY <cols> resolved via capability-matching (no query_config match).

    • Planner: asap-planner-rs/src/planner/labels.rs:15-21 (set_subpopulation_labels) — for CountMinSketch, GROUP BY columns go into aggregated_labels, and grouping_labels is left empty.
    • Engine: asap-query-engine/src/engines/simple_engine/sql.rs:318-322 (build_query_requirements_sql) always builds the capability requirement's grouping_labels from the literal SQL GROUP BY columns; asap_types/src/capability_matching.rs:82-84 (labels_compatible) requires an exact match between config.grouping_labels and requirements.grouping_labels.
    • Result: for a real planner-shaped config (grouping_labels=[], aggregated_labels=["datacenter"]), the requirement always asks for grouping_labels=["datacenter"] — mismatch every time, even though it's exactly the right aggregation.
    • Repro: built a SimpleEngine with a planner-shaped CountMinSketch(sub_type="count") config and empty query_configs, ran SELECT COUNT(cpu_usage) FROM metrics_table WHERE time BETWEEN ... GROUP BY datacenter → None. Confirmed at the find_compatible_aggregation unit level too.

    4. compatible_agg_types(Statistic::Sum) omits CountMinSketch, which is exactly what the planner emits for approximate SQL SUM (Direction B, compounds #3)

    • Planner: asap-planner-rs/src/planner/sql.rs:201-206 (get_sql_treatment_type: only MIN/MAX are Exact) → promql_utilities/src/query_logics/logics.rs:36-47 maps Sum, Approximate → AggregationType::CountMinSketch.
    • Engine: asap_types/src/capability_matching.rs:21 — Statistic::Sum => &[Sum, MultipleSum], missing CountMinSketch (contrast with Statistic::Count at line 22-25, which does list it correctly).
    • Result: a SUM query can never resolve via capability-matching regardless of labels, because its type isn't even recognized as SUM-compatible.
    • Repro: find_compatible_aggregation with a CountMinSketch(sub_type="sum") config and a Statistic::Sum requirement (matching labels) → None. Reproduced end-to-end via a plain SELECT SUM(cpu_usage) ... GROUP BY datacenter too.

    5. Planner categorically rejects nested SQL queries that the engine fully implements — CORRECTED: not an inconsistency, already fixed

    This was in the original sweep but turned out to be stale. The SQL subagent found the planner hard-rejects nested SQL (asap-planner-rs/src/planner/sql.rs:86-93, subquery depth n != 1 → Err(ControllerError::SqlParse(...))) and cited two engine tests (tests/query_equivalence_tests.rs:319-374, tests/sql_pattern_matching_tests.rs:167-196) as evidence the engine still executes this shape. On manual review after cross-referencing GitHub issues, those tests actually assert context.is_none() for exactly this query shape — the subagent misread them.

    Nested SQL matcher support was deliberately removed in commit 3e36cae (PR #504, closing issue #499, part of tracking issue #497) on 2026-07-02, predating this investigation. The engine now returns QueryError::NestedQueryUnsupported for the same shape the planner rejects. Planner and engine agree; this is not a bug. SQL still has no punting mechanism (asap-planner-rs/src/sql/generator.rs:114 always sets punted_queries: Vec::new()), and the planner's rejection still aborts the whole generate() call rather than degrading per-query — but that's a UX/robustness question, not a planner/engine inconsistency, so it's out of scope for this doc.

    On the already-known self-keyed-heap gap

    Walked generate_sql_plan (asap-planner-rs/src/sql/generator.rs:78-101): every successfully-planned SQL query — including single-level top-k — unconditionally gets a query_configs entry, so find_query_config_sql always has a match for anything the planner actually plans. The known gap (CountMinSketchWithHeap capability-matching fallback doesn't know a heap can be self-keyed, tracked separately from #498) stays dormant for planner-issued queries; it would only surface for query text the planner never saw. No concrete planner-emitted instance found.


    Cross-referenced against existing GitHub issues

    Fix priority (suggested, not decided)

    Findings #1, #3, and #4 (plus the pre-existing #501 and #267) all point at the same place: find_compatible_aggregation and its compatibility predicates (is_multi_population_value_type, compatible_agg_types, labels_compatible) have drifted from what the planner's generators actually emit, repeatedly, across multiple aggregation types. That function is worth a dedicated audit/rewrite against the planner's real output rather than continuing to patch it type-by-type. Finding #2 is separate and structural: PromQL punting has no runtime enforcement at all.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions