perf: a watch pinned to one name polls only the node that holds its entry - #48
Merged
Merged
Conversation
…ntry Every kubectl wait and kubectl delete opens a watch pinned to one name, and each of its one-second ticks re-derived the whole fleet's inventories to follow that name. While the watch holds its entry, a tick now reads only that entry's node. While it holds none, the tick sweeps the inventories until the first hit, without building the rest of the fleet. The keep-while-the-node-holds-it check is unchanged, and fleet watches are unchanged. BenchmarkClientInventoryWatchTick, informer-fed cache, median of 6 runs, arms interleaved in both orders: 26x100 2.05 ms to 0.062 ms, 26x2000 23.2 ms to 1.11 ms, 200x2000 164 ms to 1.11 ms per tick.
…uplicate A repeated Create under one name leaves two claims on two nodes. While the watch held no entry, the sweep took whichever claim answered first and filtered it afterwards, so a claim the selectors rejected could hide the one they accept and delay its Added by several ticks. The fleet poll this replaced filtered first. The sweep now skips a claim the selectors reject.
…cate they accept A repeated Create leaves one name claimed on two nodes. The pinned list resolved like Get and filtered afterwards, so a claim the selectors reject hid the one they accept, while the pinned watch already skipped it. The list now falls back to the watch's selector-aware sweep when its pick is rejected.
Contributor
Author
|
Two-host hardware A/B, 2026-09-25. Master
The record build also carried the rest of the sandbox chain (41-sbx: 28/28 PASS) with no regression. Observation, not new with these PRs: (Edited: an earlier version of this comment said the master watch reported the paused claim. It reported the running one.) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #47. Squashing #46 left #47's branch in conflict with
master, and moving it would need a force push, so this is the same commit on the mergedmaster, with the same tree.Problem
Every
kubectl waitandkubectl deleteopens a watch pinned to one name. Each one-second tick of that watch calledlistInventories: it read every node's inventory, built aSandboxfor every entry in the namespace and sorted them, only to follow one name. The cost grows with the whole fleet, and each concurrent wait pays it again. #46 listed this as its follow-up.Change
runWatchpicks its poll once: a pinned watch callspollPinned, and every other watch keepslistInventories.pollPinnedreads only that entry's node throughmatchOnNode.FirstHitand stops at the first hit. It asks no node's sandboxd, so an absent name still costs no node request.listPinnedandpollPinnedshareselectedOne, which returns zero or one sandbox that passes the selectors.docs/scaling-design.mdstates the exception to "one watcher per fleet".Behavior notes:
Cost
The claim path is unchanged. Production code is +32/−5, with no comment added. The list fallback runs only when the pick is rejected, and costs one sweep of the cached inventories.
BenchmarkClientInventoryWatchTickruns one pinned tick through the informer-fed cache that production uses. Each arm ran 6 times, interleaved, 3 rounds in each order. The figures are medians.The pinned tick grows with the entries on one node, not with the fleet. It still decodes that node's whole inventory, which is the shared decode every read path uses.
Tests
TestANamePinnedWatchPollsOnlyTheNodeThatHoldsItsEntryinpkg/scale/sandboxstore_live_test.go, under synctest:With the tick forced back to
listInventories, both fleet-enumeration assertions fail. The #46 pinned-watch tests pass unchanged.TestANamePinnedWatchSweepSkipsADuplicateItsSelectorsRejectandTestANamePinnedListSkipsADuplicateItsSelectorsRejectput a Paused claim on one node and a Running claim under the same name on another, then selectphase=Running. Each fails without its fix (8 of 8 runs) and passes with it.Gates, all with
GOWORK=off:make fmt-checkmake lint: 6 ×0 issues.go test -race -count=1 ./...asl ./...on darwin and linux: only the twopkg/envdproxyforwarder advisories (withTarget,sandboxHost), which are keptHardware verification
Pending. The next two-host round re-runs SBL-06b and SBL-08 on this branch.