Skip to content

feat(planner): read MILP accuracy from sketch-bench's saturation curves; heap from the plan - #804

Merged
zzylol merged 6 commits into
mainfrom
bump-rqe-optimizer-one-study
Oct 7, 2026
Merged

zzylol merged 6 commits into
mainfrom
bump-rqe-optimizer-one-study

Conversation

@zzylol

@zzylol zzylol commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Depends on ProjectASAP/sketch-bench #178, #179, #180, #182 and #185, all merged. Rebuilt on main after #798, #801, #803, #806 and #807.

Why

  • Accuracy came from the wrong source. The MILP took accuracy from the cost table (table_accuracy). Since sketch-bench Moved some files from asa-query-engine to asap-common #171/[Discussion] Explore better sketch storage format #178, sketch accuracy comes from the saturation study's curves, and a cost-table row is only the point it was measured at.
  • The heap was too small. sketch-bench #182 prices a top-k deployment's heap at m · k for the m windows a query merges. The emitted CMS-with-heap config still used the query's literal k, smaller than what the plan priced and what a merged answer needs.

What

  • Pin. rqe-optimizer moves to sketch-bench main (a2e0455).
  • Accuracy.
    • asap-optimizer-cli takes a required --saturation-dir, the study directory with out_grid_1e7_cost/ and out_1e9/, and passes SaturationCurves::accuracy to solve_milp.
    • solve_milp now takes the accuracy function as an argument. The tests pass table_accuracy.
    • The cost table still comes from --atomic-costs.
  • Workload facts. Each group may carry shape: {zipf_s, distinct_keys, tail_index}, fitted by the caller as in sketch-bench's small_problem. It maps to MetricFacts.data_shape.
    • It is validated like value_range: all values finite, zipf_s >= 0, distinct_keys >= 1, tail_index > 0. The error names the metric and labels.
    • A RAQE that a sketch family with cost rows could serve (a candidate under --allow-undeployable-families), on a grouping without a shape, fails before solving with MissingShape, naming the RAQE, metric and grouping. Without a shape, sketches get no accuracy. Families without rows are left to the normal Unservable path.
  • Curve loading. A failure to load the curves names the --saturation-dir path.
  • Top-k k. Each top-k query's literal k now flows into Raqe.topk_k, using sketch-bench Benchmark the precompute engine accuracy of accumulators #185's per-query k. The plan prices the heap at m · k and scales query cost by the answered k. A top-k query without a literal k fails with TopkWithoutK.
  • Heap. heapsize is the plan's own heap: the deployment's heap param, which is m · k, or the smallest measured heap that already holds it. A heap family without one is a MissingParam error.
    • This ships with the bump because the plan now prices that heap, so the emitted config has to build it.
    • The engine reads each query's k from the query itself, so the config carries no separate k.
  • Cost rows. accuracy_metric and measured_at are now required, and merge_accuracy is gone. Fixtures follow. The top-k fixtures carry rows at heap 32 and 2048, as the cost table does.

Verification

  • cargo test -p asap_planner passes: 126 unit tests plus the integration tests.
  • New tests:
    • topk_plans_with_its_own_k_and_heap_m_times_k: top-10 and top-100 plan with their own k and emit the planned heap. For k = 100 that is exactly m · k; for k = 10 it is the smallest measured heap, 32, which holds it.
    • a_shared_heap_is_the_planned_heap
    • a_heap_family_without_a_heap_is_a_missing_param
    • topk_carries_its_literal_k_and_needs_one
    • rejects_an_invalid_shape
    • a_sketch_served_grouping_without_a_shape_is_named: also covers skipping a family without rows.
    • a_group_with_a_shape_keys_the_curves
    • atomic_cost_entry_requires_accuracy_metric_and_measured_at
  • cargo clippy -p asap_planner --all-targets -D warnings and cargo fmt --check are clean.

Open

  • End-to-end run. asap-optimizer-cli against real curves and the regenerated cost table waits for the cost-table and curve rerun on sketch-bench main with asap_sketchlib 0.3.0 (sketch-bench Tracking issue for security monitoring/detection query demo #183, merged; running).
  • No cross-check. The CLI doesn't run sketch-bench's check_cost_table against the curves.

🤖 Generated with Claude Code

@zzylol
zzylol marked this pull request as ready for review October 7, 2026 18:30
…plan

Bump rqe-optimizer to sketch-bench main (a2e0455): #178 one-study cost
table, #179 KLL merge curves, #180 theory fallback, #182 top-k heap m·k.

- asap-optimizer-cli takes a required --saturation-dir and passes
  SaturationCurves::accuracy to solve_milp, which now takes the accuracy
  function instead of hard-wiring the cost table's (table_accuracy, still
  what the tests pass). The cost table stays --atomic-costs.
- Workload facts: each group may carry `shape: {zipf_s, distinct_keys,
  tail_index}`, mapped to MetricFacts.data_shape. A grouping without one has
  no sketch accuracy, so only exact accumulators serve it.
- CMS-with-heap heapsize is the heap the plan priced (the deployment's
  `heap` param, m·TOPK_K for the windows it merges), floored at the largest
  literal k it serves. It was the largest k alone, which a merged answer
  outgrows.
- Cost rows: required accuracy_metric and measured_at, no merge_accuracy;
  fixtures follow, and top-k fixtures carry heap 32 and 2048 rows as the
  table now does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol force-pushed the bump-rqe-optimizer-one-study branch from 5e6e3a4 to 17cfcd6 Compare October 7, 2026 20:17
@zzylol zzylol changed the title planner: adopt sketch-bench's curve-based accuracy and one-study cost table planner: read MILP accuracy off sketch-bench's curves; heap from the plan Oct 7, 2026
@zzylol
zzylol marked this pull request as draft October 7, 2026 20:17
- heapsize is m·k: the planned m·TOPK_K scaled by the largest literal k
  when it exceeds TOPK_K, so merged heaps still hold the top k. The plan's
  heap is required (expect), not defaulted.
- Tests assert the emitted heap equals the query's heap_needed, and check
  heap_size directly on a deployment merging m = 4 windows.
- Workload facts reject a shape that isn't finite with zipf_s >= 0,
  distinct_keys >= 1 and tail_index > 0, naming the metric and labels.
- solve_milp fails with MissingShape (RAQE, metric, grouping) before
  solving when a sketch-served RAQE's grouping has no shape.
- The CLI's curve load error names the --saturation-dir path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol zzylol changed the title planner: read MILP accuracy off sketch-bench's curves; heap from the plan feat(planner): read MILP accuracy from sketch-bench's saturation curves; heap from the plan Oct 7, 2026
…ked heap

- heap_size takes m from the plan, the most windows a served query's
  lookback merges (query_instance_count), and emits max(planned heap,
  m · max(TOPK_K, largest literal k)); the row's heap can be a measured size
  above m · TOPK_K. No division by TOPK_K, checked m · k (HeapSize error),
  and a row without a heap is MissingParam, not a panic.
- require_shapes only fires for a sketch family that has cost rows and is a
  candidate; others are left to the Unservable path.
- GroupEntry's doc states the MissingShape behavior; DuplicateLabels is
  checked before shape validation; DataShape is built from the locals.
- Tests: m from the lookback under a larger measured heap, overflow and a
  missing heap, and require_shapes skipping a family without rows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol marked this pull request as ready for review October 7, 2026 20:26
…tchlib 0.3.0

sketch-bench #183 moved to asap_sketchlib 0.3.0, where the top-k heap's
update is O(log k) and its memory counts the heap index.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
)

Bump rqe-optimizer to sketch-bench main 30e725a4, where top-k's k is a
per-query variable: the heap needed is m·k and query cost scales by the
answered k.

- item_to_raqe sets Raqe.topk_k to the query's literal k (the largest, should
  an item carry several); a top-k query without one is MilpError::TopkWithoutK.
- heapsize is the plan's own heap, the deployment's `heap` param, with no
  output-side scaling: the plan already sized it for m·k. A heap family
  without one is still MissingParam. The engine reads each query's k from
  the query, so the config carries no separate k.
- Tests: top-10 and top-100 plan with their own k and emit the planned heap
  (m·k, or the smallest measured heap that holds it); a shared heap is the
  planned one; the literal k reaches the Raqe.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol merged commit 847bf1e into main Oct 7, 2026
8 checks passed
@zzylol
zzylol deleted the bump-rqe-optimizer-one-study branch October 7, 2026 22:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants