Repository navigation
Conversation
zzylol
force-pushed
the
stack/univmon-l2-accuracy
branch
from
October 5, 2026 03:07
18cb590 to
ba716f2
Compare
zzylol
force-pushed
the
stack/satisfies-error-metric
branch
from
October 5, 2026 03:07
ba0c9e2 to
64ebe46
Compare
This was referenced Oct 5, 2026
Stage 3 compared only a summary guarantee's bound and failure probability with the target, so a bound in one metric (Count-Min's L1 frequency error, say) could satisfy a target on another statistic (a distinct count). `AccuracyModel::answers` maps each statistic to the metrics its epsilon is stated in, and `build_violation` checks it before `satisfies`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
force-pushed
the
stack/univmon-l2-accuracy
branch
from
October 5, 2026 06:21
ba716f2 to
e34c85c
Compare
zzylol
force-pushed
the
stack/satisfies-error-metric
branch
from
October 5, 2026 06:21
64ebe46 to
4f65ff1
Compare
This was referenced Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: #574 → #620 → #618 → #621 → #627 → #625 → #628 → #632 → #634 → #616 → #617 → #622 → #624 → #629 → #630 → #631 → #633 → #635 → #636 → #637
Problem
DefaultAccuracyModel::satisfiescompares only a guarantee's bound and δ with the target. Stage 3'sbuild_violationtherefore accepts a bound stated in the wrong metric. For example, a Count-Min L1 frequency bound (Frequency) passes a target on a distinct count (Cardinality) whenever its number is small enough.Changes
AccuracyModel::answers(statistic, guarantee): a new trait method with a default implementation. It is a separate method becausesatisfies(guarantee, target)never sees the statistic. Its callers in Pass 1 check composed root guarantees, where noSketchStatisticexists.An exact guarantee answers every statistic.
Otherwise the guarantee's metric must be one the statistic's ε is stated in:
Rank,RelativeValueCardinality,RelativeValueRelativeValueFrequency,L2FrequencyFrequency,L2Frequency,TopKMembership(the score bound Pass 1 already checks top-k targets against)A deployment with a registered cross-metric conversion can override
answers.EstimatorAccuracydelegates it to its base model.build_violationchecksanswersbeforesatisfies. On a mismatch it rejects with "… guarantees a bound, which does not bound ".answers_requires_the_statistics_metric.a_bound_in_another_metric_misses_the_target: a CountSketch+heap build passes for its TopK readout, but is rejected for Cardinality even at ε = δ = 0.5.No example's selection changes. Examples planner-layering-1, 2, 3a, 3b, 4a and 4b select the same plans as before, and none of their rejections has the new reason.
Test plan
cargo fmt --all --checkcargo clippy --workspace --all-targets -- -D warningscargo test --workspace: 1679 passed, 0 failed, 21 ignored (re-run after the rebase on main d4869a7)Stacked on #621.
🤖 Generated with Claude Code