refactor(runtime)!: strategy evolution on the search kernel - #1441
Merged
Merged
Conversation
runStrategyEvolution now runs Eval's runSearch with beam({ width: 2 }) and
asha() by default. Every strategy is a node (the root, or an authored
module content-addressed by its source); the author reads train results
and its parent's source; the claim runs once on the sealed test split.
The ledger is the only checkpoint and the report is its projection.
Deleted: the generation loop, its JSON checkpoint and archive, the
champion helpers, promotionGate, and the unit tests of the touched
modules. runAgentic shots and analyst calls now report billed dollars so
evolution cells can carry complete cost accounting.
An authoring an interrupted process lost has an unknown cost and holds the ship decision; the report now says how many.
Adds examples/strategy-evolution (offline: a deterministic worker and a scripted author; a cli-bridge author when BRIDGE_URL is set), fixes the counter examples, whose worker profile never enabled its tools, and replaces promotionGate and the old evolution in the curated docs.
tangletools
previously approved these changes
Sep 27, 2026
tangletools
left a comment
Contributor
There was a problem hiding this comment.
✅ Auto-approved PR — 8aff2056
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-27T10:42:52Z
Each runAgentic run is one paid call on the search cost ledger through Runtime's searchExecutor on a dedicated in-process lane: finished attempts are recorded under their run id and adopted after a restart, each attempt is traced with agent.branch.id, and workerSlots bounds cells across searches. One executor port for every cell Runtime runs.
tangletools
approved these changes
Sep 27, 2026
tangletools
left a comment
Contributor
There was a problem hiding this comment.
✅ Auto-approved PR — 912610d8
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-27T10:50:53Z
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Package R5 of the search-tree design (
final-design.md§3 runtime row, §6.8, §10 R5): strategy evolution runs on Eval's search kernel withbeamandasha, and the archive is deleted.What changed
runStrategyEvolutionrunsrunSearch(agent-eval 0.198) withbeam({ width: 2 })+asha()by default. Every strategy is a node: theroot(defaultrefine) or an authored module content-addressed by the sha256 of its source. The author (the kernel's proposer) reads the strategy contract, the task tools, the train summary, the train tournament and its parent's source; ranking uses the private selection split; the claim runs the root and at most 3 finalists together on the sealed test split, once, with the power check and Bonferroni. The ledger at<outDir>/<searchId>/ledger.jsonlis the only checkpoint.EvolutionReportis a projection of the ledger: strategies (lineage, status, source, gzip bits, refusal, selection estimate vs root), the claim and its re-derivation, Runtime's ship decision, per-split tournaments and spend.strategyTournament(state, split, names)is exported.EvolutionCheckpointand its archive,EvolutionArchiveNode,EvolutionCandidate,EvolutionGeneration,EvolutionBandInfo,ReproductionCheck,ChampionPick,ChampionPolicy,pickChampion,selectChampion,discriminatingMeans, andpromotionGate(its only caller was the old loop). Unit tests of the touched modules are deleted:tests/kernel/strategy-evolution.test.ts,tests/kernel/strategy-suite.test.ts,src/runtime/run-benchmark.test.ts.searchExecutoron onededicatedLaneofconcurrencyslots: eachrunAgenticrun is one paid call on the search's cost ledger, finished attempts are recorded under their run id and adopted after a restart, every attempt is a trace whose root span carriesagent.branch.id, andworkerSlotsbounds cells across searches. A run a killed process left pending settles with an unknown cost.authorStrategysplits intorequestStrategySource(a failed call throws; a reply without a module iscode: null) andloadAuthoredStrategy(lint, write at a content-addressed path, import).runAgenticshots and analyst calls now carry provider-billed dollars. Before this, every shot reportedusd: 0, usdKnown: false, so no evolution cell could ever have complete cost accounting and Runtime's ship rule would hold every search.Observation.usage.costUsdis additive.examples/strategy-evolution(new, offline): deterministic worker and a scripted author; a cli-bridge author whenBRIDGE_URLis set. The counter examples' worker profile never enabled its tools, soexamples/strategy-suitescored 0 % on main; fixed.bench/src/swe-self-improve.mts, curated docs and generated API pages follow. Release note:.release-notes/strategy-evolution-on-search.md(minor, breaking, with migration).Deviations from the design (build note appended to
final-design.md)root, not a gen0 field of baselines: the kernel has one root and the claim tests finalists against it. Pick the best baseline withrunBenchmarkfirst.cost(non-inferiority) objective, band screening, the reproducer check,lossesDetail: 'binary',onPhase/onTaskare deleted: the kernel's claim is the one decision, and the binary channel's leakage bound existed because the old loop picked champions on the tasks the author saw.SearchStateViewcells (score, outcome, known-plus-floor dollars), notBenchmarkCell: the state carries no progression, wall time or tokens.open,tools,scoreorcloseduring an attempt iserroredand retried; a strategy that throws isfailedat score 0; anything elsepassed.invalidnode with its reason.runtimeShipDecision, now shared withsearchMethod) withminimumLift0. A cell or authoring a killed process interrupted has an unknown cost, so a resumed search keeps its verifiedshipclaim but holds the decision; the report countsunknownCostCellsandunknownCostOperations.Proof (real runs, no unit tests, no model spend)
On drew-gtr-pro, built package,
_mq-notes/R5-proof/:Example output:
carry-forwardships at test lift 0.276 [0.195, 0.351] on 24 tasks;agent-eval search show <ledger>prints the search.Kill-resume (
kill-resume.sh <outDir> <seed> 10, SIGKILL of the exact node pid 0.9 to 2.9 s after each start, then resume):A killed run may have spent money, so its cell settles with an unknown cost and Runtime's ship rule holds; nothing about the claim changes.
Faults (
fault-run.mts: every 7thopenthrows; the first authored module always throws): 7environment-faultattempts settled as retried errors (attempts 2 and 3 present), the throwing strategy's 8 cellsfailedat 0 and pruned at selection Δ −0.546, 0 cells allocated twice, 0 attempt gaps.Gate on drew-gtr-pro after merging main (R3 #1437, R4 #1436, 0.281.0):
pnpm install --frozen-lockfile,tsc --noEmit,tsc -p tsconfig.examples.json,pnpm build,pnpm docs:api,pnpm docs:freshness(clean). Fullvitest runon the branch before the merge: 346 files passed, 1 skipped.Not run: a live author through cli-bridge. The gtr
claude-codeseat's OAuth expired ("OAuth session expired and could not be refreshed");opencodeandkimi-codebridge turns do not report their served model, whichprofileChatClientrefuses;pirequires an fs-jail. The authoring path throughprofileChatClientis the oneauthorStrategyalready used.