Skip to content

refactor(runtime)!: strategy evolution on the search kernel - #1441

Merged
drewstone merged 6 commits into
mainfrom
refactor/strategy-evolution-on-search
Sep 27, 2026
Merged

drewstone merged 6 commits into
mainfrom
refactor/strategy-evolution-on-search

Conversation

@drewstone

@drewstone drewstone commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Package R5 of the search-tree design (final-design.md §3 runtime row, §6.8, §10 R5): strategy evolution runs on Eval's search kernel with beam and asha, and the archive is deleted.

What changed

  • runStrategyEvolution runs runSearch (agent-eval 0.198) with beam({ width: 2 }) + asha() by default. Every strategy is a node: the root (default refine) or an authored module content-addressed by the sha256 of its source. The author (the kernel's proposer) reads the strategy contract, the task tools, the train summary, the train tournament and its parent's source; ranking uses the private selection split; the claim runs the root and at most 3 finalists together on the sealed test split, once, with the power check and Bonferroni. The ledger at <outDir>/<searchId>/ledger.jsonl is the only checkpoint.
  • EvolutionReport is a projection of the ledger: strategies (lineage, status, source, gzip bits, refusal, selection estimate vs root), the claim and its re-derivation, Runtime's ship decision, per-split tournaments and spend. strategyTournament(state, split, names) is exported.
  • Deleted: the generation loop, EvolutionCheckpoint and its archive, EvolutionArchiveNode, EvolutionCandidate, EvolutionGeneration, EvolutionBandInfo, ReproductionCheck, ChampionPick, ChampionPolicy, pickChampion, selectChampion, discriminatingMeans, and promotionGate (its only caller was the old loop). Unit tests of the touched modules are deleted: tests/kernel/strategy-evolution.test.ts, tests/kernel/strategy-suite.test.ts, src/runtime/run-benchmark.test.ts.
  • Cells run through R3's searchExecutor on one dedicatedLane of concurrency slots: each runAgentic run is one paid call on the search's cost ledger, finished attempts are recorded under their run id and adopted after a restart, every attempt is a trace whose root span carries agent.branch.id, and workerSlots bounds cells across searches. A run a killed process left pending settles with an unknown cost.
  • authorStrategy splits into requestStrategySource (a failed call throws; a reply without a module is code: null) and loadAuthoredStrategy (lint, write at a content-addressed path, import).
  • runAgentic shots and analyst calls now carry provider-billed dollars. Before this, every shot reported usd: 0, usdKnown: false, so no evolution cell could ever have complete cost accounting and Runtime's ship rule would hold every search. Observation.usage.costUsd is additive.
  • examples/strategy-evolution (new, offline): deterministic worker and a scripted author; a cli-bridge author when BRIDGE_URL is set. The counter examples' worker profile never enabled its tools, so examples/strategy-suite scored 0 % on main; fixed.
  • bench/src/swe-self-improve.mts, curated docs and generated API pages follow. Release note: .release-notes/strategy-evolution-on-search.md (minor, breaking, with migration).

Deviations from the design (build note appended to final-design.md)

  1. One root, not a gen0 field of baselines: the kernel has one root and the claim tests finalists against it. Pick the best baseline with runBenchmark first.
  2. The cost (non-inferiority) objective, band screening, the reproducer check, lossesDetail: 'binary', onPhase/onTask are deleted: the kernel's claim is the one decision, and the binary channel's leakage bound existed because the old loop picked champions on the tasks the author saw.
  3. The tournament projects SearchStateView cells (score, outcome, known-plus-floor dollars), not BenchmarkCell: the state carries no progression, wall time or tokens.
  4. Cell outcomes: an environment throw from open, tools, score or close during an attempt is errored and retried; a strategy that throws is failed at score 0; anything else passed.
  5. An author transport failure pauses the search (the same call resumes); a reply without a module, a lint refusal or a module with no default Strategy is an invalid node with its reason.
  6. The ship rule is Runtime's (runtimeShipDecision, now shared with searchMethod) with minimumLift 0. A cell or authoring a killed process interrupted has an unknown cost, so a resumed search keeps its verified ship claim but holds the decision; the report counts unknownCostCells and unknownCostOperations.

Proof (real runs, no unit tests, no model spend)

On drew-gtr-pro, built package, _mq-notes/R5-proof/:

$ node --import tsx .evolution-proof/run.mts .evolution-u2 .evolution-u2/report-1.json
{"searchId":"evolution-688a8356c79be8e56dda9524","closeReason":"max-nodes","strategies":["sample:rejected","one-shot:pruned","carry-forward:selected","refused-module:invalid","refused-module~4:invalid","carry-forward-steered:pruned"],"selected":"carry-forward","claim":"ship","verification":"verified","decision":"ship","spend":{"knownUsd":0.00539,"floorUsd":0,"unknownCostCells":0,"unknownCostOperations":0},"ms":42909}
$ node --import tsx .evolution-proof/run.mts .evolution-u2 .evolution-u2/report-2.json   # same inputs, closed ledger
{... identical ..., "ms":560}
reports byte-identical
$ node --import tsx .evolution-proof/audit.mts <searchDir>
{"entries":218,"cellsAllocated":92,"allocatedTwice":0,"cellsSettled":92,"attemptsSettled":92,"settledTwice":0,"attemptGaps":0,"operations":6,"operationsUnrecorded":0,"operationsInterrupted":0}
$ node --import tsx .evolution-proof/project.mts <searchDir> report-1.json   # fresh process, ledger bytes only
{"ledgerLines":218,"closed":"max-nodes","nodes":6,"edges":{"explicit":6,...},"cells":{"allocated":92,"settled":92,"cancelled":0,"open":0},"tournamentRows":{"train":2,"selection":12,"test":24},"tournamentEqualsReport":{"train":true,"selection":true,"test":true},"claimEqualsReport":true,"claimVerification":{"status":"verified"},"namesEqualReport":true,"statusEqualsReport":true,"cellsWithSameIdTwice":0}
$ wc -l spans.otlp.jsonl; ls attempts | wc -l
92 spans.otlp.jsonl
92

Example output: carry-forward ships at test lift 0.276 [0.195, 0.351] on 24 tasks; agent-eval search show <ledger> prints the search.

Kill-resume (kill-resume.sh <outDir> <seed> 10, SIGKILL of the exact node pid 0.9 to 2.9 s after each start, then resume):

Seed Kills Final nodes/cells allocatedTwice / settledTwice / gaps Scores, outcomes, strategies, claim vs uninterrupted Killed runs reconciled as unknown cost Decision
1 10 6 / 92 0 / 0 / 0 equal; ship verified 6 cells hold
2 10 6 / 92 0 / 0 / 0 equal; ship verified 11 cells hold
3 10 6 / 92 0 / 0 / 0 equal; ship verified 5 cells + 1 authoring hold

A killed run may have spent money, so its cell settles with an unknown cost and Runtime's ship rule holds; nothing about the claim changes.

Faults (fault-run.mts: every 7th open throws; the first authored module always throws): 7 environment-fault attempts settled as retried errors (attempts 2 and 3 present), the throwing strategy's 8 cells failed at 0 and pruned at selection Δ −0.546, 0 cells allocated twice, 0 attempt gaps.

Gate on drew-gtr-pro after merging main (R3 #1437, R4 #1436, 0.281.0): pnpm install --frozen-lockfile, tsc --noEmit, tsc -p tsconfig.examples.json, pnpm build, pnpm docs:api, pnpm docs:freshness (clean). Full vitest run on the branch before the merge: 346 files passed, 1 skipped.

Not run: a live author through cli-bridge. The gtr claude-code seat's OAuth expired ("OAuth session expired and could not be refreshed"); opencode and kimi-code bridge turns do not report their served model, which profileChatClient refuses; pi requires an fs-jail. The authoring path through profileChatClient is the one authorStrategy already used.

runStrategyEvolution now runs Eval's runSearch with beam({ width: 2 }) and
asha() by default. Every strategy is a node (the root, or an authored
module content-addressed by its source); the author reads train results
and its parent's source; the claim runs once on the sealed test split.
The ledger is the only checkpoint and the report is its projection.

Deleted: the generation loop, its JSON checkpoint and archive, the
champion helpers, promotionGate, and the unit tests of the touched
modules. runAgentic shots and analyst calls now report billed dollars so
evolution cells can carry complete cost accounting.
An authoring an interrupted process lost has an unknown cost and holds the ship decision; the report now says how many.
Adds examples/strategy-evolution (offline: a deterministic worker and a scripted author; a cli-bridge author when BRIDGE_URL is set), fixes the counter examples, whose worker profile never enabled its tools, and replaces promotionGate and the old evolution in the curated docs.
tangletools
tangletools previously approved these changes Sep 27, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 8aff2056

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-27T10:42:52Z

Each runAgentic run is one paid call on the search cost ledger through Runtime's searchExecutor on a dedicated in-process lane: finished attempts are recorded under their run id and adopted after a restart, each attempt is traced with agent.branch.id, and workerSlots bounds cells across searches. One executor port for every cell Runtime runs.

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 912610d8

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-27T10:50:53Z

@drewstone
drewstone merged commit b362f61 into main Sep 27, 2026
4 checks passed
@drewstone
drewstone deleted the refactor/strategy-evolution-on-search branch September 27, 2026 14:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants