Skip to content

refactor(memory)!: remove the unused memory experiment, improvement, and benchmark runners - #241

Merged
drewstone merged 1 commit into
mainfrom
refactor/remove-memory-experiments-20261006
Oct 6, 2026
Merged

drewstone merged 1 commit into
mainfrom
refactor/remove-memory-experiments-20261006

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Problem

/memory and /benchmarks carried a second experiment and improvement engine: matched-arm learning experiments, runAgentMemoryImprovement with run leases, cost meters, and activation, and a benchmark suite with industry smoke catalogs, qrels import, scoring, and memory recovery runners. That is 9,545 source lines and 5,746 test lines. Nothing outside this repository imports any of it. Agent Eval already owns shared experiment and improvement machinery (AGENTS.md: "shared evaluation and improvement machinery belongs to agent-eval").

Change

  • Delete src/memory/experiment/, src/memory/improvement/, their barrels, memory/run-control.ts, memory/attempt-log.ts, candidate-ranking.ts, and every src/benchmarks/ module except adapters.ts.
  • Delete their tests: tests/memory/experiment-*, tests/memory/improvement.test.ts, tests/benchmarks/, the support helpers, and the improvement-candidate cases in the Mem0 and Graphiti tests.
  • /benchmarks now exports only createInMemoryBenchmarkAdapter and createNoopMemoryBenchmarkAdapter. Supervisor Lab imports both.
  • The README memory and benchmark sections, the package-boundary table, and the skill drop the removed runners. verify-package.mjs checks createAgentMemoryBranch on /memory in place of runAgentMemoryImprovement.
  • Add to the unpublished 20.0.0 changelog entry (210 exports removed).

Kept: memory adapters (Mem0, Neo4j, Graphiti, Hindsight), branches, retrieval holdout, runBoundedMemoryLifecycle and its errors (Hindsight uses them), play memory tools, schemas, and source records.

Consumer evidence

The same scan as #240: identifiers imported from the package, including /memory and /benchmarks subpaths and dynamic imports. It covered the Mac ~/webb and ~/company trees, ~/code on beelink1 and drew-gtr-pro, and fresh default-branch clones of the 22 repositories that depend on the package, plus discovery, agent-eval, agent-sdk, and braid. None of the removed names is imported. /memory imports in use: Agent Runtime (AgentMemoryAdapter, branch and snapshot APIs, createPlayMemoryTools, defaultGetMemoryContext), Blueprint Agent and Supervisor Lab (memory types, holdout, renderMemoryContext). /benchmarks imports in use: Supervisor Lab's archive/bench/memory (the two adapters). All of these stay.

Why this is the right long-term shape

Knowledge keeps the memory contract and provider adapters. Experiments, comparison, and promotion decisions run in Agent Eval, which has one maintained implementation.

Cost

69 files, 11 insertions, 15,568 deletions. This rides the unpublished 20.0.0 major from #240 (latest published is 19.1.5), so it adds no second major. Risk: an unlisted private caller of the removed runners. Rollback: revert, or stay on 19.x.

Verification (beelink1, merged with origin/main)

  • pnpm install --frozen-lockfile, pnpm lint (only the existing proposals.ts warning), pnpm typecheck (source and contracts), pnpm build, and pnpm run api:surface (740 exports) all pass.
  • pnpm test with network tests: 72 files, 754 passed, 2 skipped, 0 failed.
  • node scripts/check-version-bump.mjs passes: 210 export changes are paid for by 19.1.5 -> 20.0.0.
  • Every relative markdown link resolves, and git merge-tree --write-tree origin/main HEAD is clean.
  • verify:package is still blocked upstream for main and this branch alike: Agent Core 0.10.3 requires Interface 3 (see refactor(research)!: remove the unused web research drivers and TCloud #240).

…and benchmark runners

Nothing outside this repository imports runAgentMemoryExperiment,
runAgentMemoryLearningExperiment, runAgentMemoryImprovement,
runKnowledgeBenchmarkSuite, or runMemoryAdapterBenchmark. They duplicated
experiment and improvement machinery that Agent Eval owns. Memory adapters,
branches, holdout, lifecycle bounds, play memory tools, and the benchmark
adapters Supervisor Lab imports stay.

BREAKING CHANGE: the memory experiment, memory improvement, and benchmark
suite exports are removed.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 4d4c0650

Blanket team auto-approval is intentional. This is not a code review.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-10-06T07:10:47Z

@drewstone
drewstone merged commit 8525846 into main Oct 6, 2026
@drewstone
drewstone deleted the refactor/remove-memory-experiments-20261006 branch October 6, 2026 07:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants