Skip to content

feat(dataset-analysis): 6h and 24h windows over a day of Alibaba - #812

Open
zzylol wants to merge 2 commits into
mainfrom
dataset-analysis-long-windows
Open

zzylol wants to merge 2 commits into
mainfrom
dataset-analysis-long-windows

Conversation

@zzylol

@zzylol zzylol commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

What

The AutoSketch vs. planner evaluation (#777) needs trace RQEs with 6h and 24h lookbacks. This PR refits the dataset analysis with them.

  • Queries: every range query in queries/alibaba_v2022.yaml and queries/google_2011.yaml also fits 6h and 24h.
  • Data: fetch_data.sh fetches the first day of Alibaba CallGraph and MCRRTUpdate (480 three-minute shards; was 120, i.e. 6 hours), so a 24h window fits. MSMetrics and NodeMetrics already covered a day; Google already had about 7 days. About 207 GB in all.
  • results/skew_summary.csv refit:
    • Google's 48 earlier rows are byte-identical; 32 rows are new.
    • Alibaba has 115 rows (was 69). Earlier ranges change, because each window now has about 4× the evaluations: for example mcr_by_msname 1h, n_evals 301 → 1381, lower θ 1.39 → 0.86.
    • BOOM rows are unchanged.

Two fixes the larger run needed

  1. Memory in resolve_boundaries. With 3-minute files almost every sample is a file's first or last, and a day of MSRTMCR (3.9e9 rows) peaked at 220 GB on a 251 GB machine, then was killed. It now stitches in chunks of 60 steps. A sample counts for at most lookback steps, so a chunk owning steps [start, end) reads the boundary samples up to end + lookback and the aggregates are the same as one pass (test_boundary_chunks_match_one_pass). Peak RSS is now 98 GB.
  2. merge_samples above 1e9 values. numpy's multivariate_hypergeometric refuses a total of 1e9 or more, which a day of Alibaba rt values exceeds. At or above that limit each part's share is drawn multinomially, capped at its prefix. Drawing 1e5 of ≥ 1e9 values, with and without replacement agree (test_merge_samples_over_a_billion). Below the limit nothing changes.

Run

On one 56-core, 251 GB CloudLab node each, 48 workers:

  • Google: 53 min, main process peak 45 GB.
  • Alibaba: 10.5 h, 98 GB. CallGraph's 6h windows over high-cardinality keys dominate.

Checks

  • 44 unit tests pass, including the two new ones.
  • black, flake8 and mypy pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

zzylol and others added 2 commits October 8, 2026 15:16
- Every range query also fits 6h and 24h windows (Alibaba and Google).
- fetch_data.sh fetches the first day of Alibaba CallGraph and MCRRTUpdate
  (480 shards; was 6 hours), so a 24h window fits; MSMetrics and
  NodeMetrics were already a day. About 207 GB in all.
- resolve_boundaries stitches instant samples in chunks of 60 steps. With
  3-minute files almost every sample is a boundary one, and a day of
  MSRTMCR peaked at 220 GB in one pass; a sample counts for at most
  lookback steps, so each chunk reads its steps plus the lookback and the
  aggregates are the same (test). Peak is now 98 GB.
- merge_samples draws its shares multinomially when the total reaches
  1e9, numpy's limit for multivariate_hypergeometric (a day of Alibaba rt
  values); drawing 1e5 of 1e9 or more, the two agree (test).
- results/skew_summary.csv refit: Google's 48 earlier rows are unchanged;
  Alibaba's change with 4x the evaluations per window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
… fit

NodeMetricsUpdate's first day spans steps 1..1439, one short of a 24h
window, so node_cpu_by_nodeid and node_cpu_p99 had no 24h evaluation.
Three 12-hour shards (steps 1..2159) give each 720; NodeMetrics is refit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 8, 2026
- §4: plans priced by use, w_cpu * AUC(CPU) + w_mem * AUC(memory), CPU
  elastic: ingest on ceil(rho) workers split by sample, compaction at
  window close, query jobs; memory as ingest, storage, compaction and
  query, each counted once; latency is a batch's longest chain (model in
  sketch-bench #188). A peak-billed cost model was dropped.
- §5: latency is reported, in two versions: no SLA, with each method's
  cost-latency frontier, and a batch latency SLA over {100 ms .. 10 s}.
- §3: PerQuery is the same MILP over each RQE's own candidates (no sharing).
- §6: the evaluated set is the mixed template set, with shared r in {1, 8}
  and metrics m in {1, 8, 16}; the dashboard set is not evaluated.
- §7: AutoSketch's planning time is search plus its measured benchmark
  (approxbench accuracy runs at 1e8 items); figures are the frontier, cost
  vs. SLA and planning time; absolute costs only.
- §8-§10: PRs (#188, ASAPQuery #812, #138 stacked on #188), decisions Q3,
  Q5-Q7, limitations (elastic CPU, compaction merge cost).
- §11: results on the synthetic mixed set.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant