Skip to content

Roadmap: ship Owen Alpha and complete the Rust production vertical after #214 #250

Description

@PhysShell

Status

Master roadmap after the completion of P-022 step 4 in #214 / PR #249.

Reconciled at 206e9c71ec62ceef2668e74d1c62228a402a03a2 (PR #347 merge — #261's 261.B, the production Rust OwnIR executable own-cli ownir, landed: the own-cli crate with its single ownir subcommand, a Python-authored CLI fixture replayed with zero Python on Linux and Windows CI, both failure-mode rulings measured under an off-by-default fault-injection feature, and the ownir_version Version family fixed to byte parity — V1/V2/V4 declared and V3 reproduced per #262's parser-domain rulings — carried onto P-022 row 7b, the proposals index and rust/README.md; #261 subsequently closed completed). Reconciled at 48b8799ee54b6152161d4b90a20d03de41170f5f (PR #346 merge — the #261 decision-packet ratification of 2026-09-08, C-1..C-5 with the exit-code rulings and the owner's rulings on the four readings, carried onto P-022 rows 7b and 8, the proposals index and rust/README.md; this body, #261, #262 and the new #345 moved with it), at f48780630780b7f38f239469a1d7d0a87f30d356 (PR #344 merge, the CI re-record of #260's sweep) and at 4520a543e0c47a886217d065b3d80920199d4f93 (PR #343 merge), carrying #260's final acceptance — the sweep onto this surface together with docs/proposals/P-022-rust-core-migration.md and the proposals index. Previous reconciliations: b05b38a (PR #342, #260's acceptance surfaces over the committed corpus), 21fb0c3 (PR #341, #259 final acceptance), 834f295 (PR #340, which also carried #337, #338 and #339), 3fc1246 (PR #336), 0738d29 (PR #325), 984de7d (PR #324) and fdcb222 (PR #322). Statuses are checkpoint-level: each step names its completed checkpoints, its remaining acceptance, its normative blocker (what its acceptance actually requires) and this roadmap's preferred sequencing (what order is cheapest). Conflating the last two is what made the P-022 table drift twice.

How to keep parity work from drifting again is now written down: P-022 § Parity-work discipline. It is the single home for those rules — deliberately not copied elsewhere, because two copies of one law drift.

Since 3fc1246 no count is typed on a status surface: the census, the surface inventory and every mutation campaign are rendered from the tree into docs/generated/ by scripts/render_checkpoint_status.py, and the Python suite fails while a fragment is stale, a campaign result no longer matches its definition, was taken on a dirty tree, missed a required catcher, names a catcher that no longer exists, or names a commit the tree does not descend from.

Two independent outcomes must now be delivered:

  1. Owen Alpha is actually published and verified from clean external consumers.
  2. Rust participates in the real C# → OwnIR → verdict path, first in shadow mode and then through an explicit cutover gate.

The release track must not wait for the full Rust migration. The Rust track must not use the release as permission to weaken parity.

Current baseline

Completed:

Still missing:

Explicitly not missing, because it was struck rather than deferred: a Rust .ownreport.json. See the #256 entry below.

Non-negotiable migration rules

  • Python remains the oracle until a separate cutover decision.
  • A Rust/Python divergence is a Rust bug unless behavior changes in a separate Python-first PR.
  • Migration PRs do not add diagnostic rules, change severity, broaden Roslyn heuristics, or weaken fixtures.
  • Every layer owns a frozen Python-authored parity surface; steady-state Rust tests run with zero Python.
  • Production crate edges remain CI-enforced.
  • own-codegen remains independent of own-analysis and own-diagnostics.
  • own-bridge feeds facts and maps verdicts; it must not duplicate analysis algorithms.
  • Owen Alpha publication does not depend on Rust-default cutover.
  • Parity work follows P-022 § Parity-work discipline — oracle over reviewer prose, mutation over plausible tests, no fail-fast during mutation campaigns, insertion-stable generated goldens.

Child issues and execution order

A. Owen Alpha release

Release DAG:

#252
 ├─> #253
 └─> #254

The release track is independent of the Rust-default cutover.

B. P-022 Rust production vertical

Each entry states its normative blocker first; preferred ordering is marked as such and is advisory.

Rust DAG (unchanged — normative dependencies only):

#251
 ├─> #255 ─> #256 ───────────────┐
 ├─> #257                        │
 └─> #258 ─> #259 ─> #260 ───────┼─> #262
                    └─> #261 ────┘

#345, the residual .own/dev CLI split out of #261, hangs off #257 for its emit slice and is not on the #262 path; the figure above is unchanged because it draws the cutover's normative dependencies only.

Preferred queue: #262 is next. #261's 261.B production OwnIR executable landed (PR #347, 206e9c7); #261 is closed completed. #260 is closed at final acceptance and off this queue, with cp5, 4b and the coordinate-domain decision. In parallel and off the critical chain: #257 (without it #345's emit slice cannot close), #263 (without it #262's performance gates have no baseline), #345 now that #261's own-cli skeleton exists, and #269's reconciliation as independent cleanup.

The defensive limits that used to head this queue landed in #326, and their position was load-bearing rather than tidy: they changed what the reference accepts, so they had to land Python-first and cp1 had to be re-measured against them rather than merged beside them. The coordinate-domain decision was the same kind of item and landed the same way (PR #341); #260's two decisions were decided the same way and landed in PR #342.

C. IDE/incremental path

IDE DAG:

#263 ──────────────┬─> #264
                  └─> #265
#255/#256 ───────────> #264
#259/#260 ───────────> #265

D. P-034

Production implementation remains blocked until the runtime marker/helper/escape-hatch contract is finalized. Do not duplicate call-site use-after-dispose rules.

Recommended agent allocation

Owner decision (not delegable)

#269 reconciliation: what is delivered (trace, first-divergence reduction) and what is deferred (the bounded minimizer, a named explain-divergence command)
BR-V5 on the protocol path (spec sentence vs `_protocol_findings`; Python-first, blocks nothing)

#260's decisions are resolved (D-4..D-7, B-2, B-3, R-1, R-2), #261's packet is ratified (C-1..C-5, 2026-09-08) and its 261.B implementation landed (PR #347, 206e9c7); #269's reconciliation heads this list now.

Local/corpus-capable agent

#263 measurements

Strong agent

#345: the residual subcommands on the same `own-cli` binary that #261's 261.B built (PR #347); `emit` only after #257

#255, #256, #258, #259 and now #261's 261.B (PR #347) are complete and drop out of this chain. #259 is closed at final acceptance — do not re-open it or any of its checkpoints; the declared OD-1 boundary is on the exclusion ledger, not open work.

Separate strong or medium-strong agent

#257

Keep codegen isolated from analysis.

Medium agent

#252 verification packet
#253/#254 release evidence and external checks

Status-drift rule

A step is never described by a single Implemented/Missing bit. Any status edit to this issue or to docs/proposals/P-022-rust-core-migration.md must state, per step: completed checkpoints, remaining acceptance, normative blocker, and preferred sequencing. Both surfaces are updated in the same change — a reconciliation that touches only one of them replaces a stale pair with a contradictory pair. The proposals index row counts as a third surface for the same fact.

A child issue's own body is a fourth surface when its acceptance turns out to be wrong. #256 is the worked example: its requirements described a .ownreport.json the project does not have and had already refused to build, so the correction belongs in the issue, in this roadmap, in P-022 and in the index — together, or not at all. #259 is the second: its checkpoint list had no row for the obligation-protocol analysis while its final acceptance required it, so 4b was added to the child issue, to P-022 and here in the same move.

A child issue's open/closed bit is a fifth surface, and #259 is its worked example too: GitHub closed it when PR #339 merged, because its parser read the body's "does not close #259" as close #259 — the negation is not parsed — while the sentence meant the opposite and this queue still read "→ #259 final acceptance". A closed issue with an open acceptance is a contradictory pair with the roadmap. It was reopened, and closed by hand at the #341 merge once the acceptance was reached. Checkpoint PRs reference their issue as Refs #N and never put a closing keyword before an issue number, negated or not; nothing but the final-acceptance change closes it, and the owner does that by hand together with the body update.

Acceptance-evidence surface (parity checkpoints only)

Parity checkpoints carry one more surface, and it is not a status surface: the frozen ledger. A checkpoint whose acceptance is "two implementations agree" is proved by an artifact that can share the implementation's blind spots, and a green matrix over an incomplete ledger is indistinguishable from a green matrix over a complete one. #259 cp1 is the worked example, three times over — 0/0/0 over 77 controls, then 58 permissive documents and 9 category mismatches once the ledger was rebuilt from the reference instead of from the author's reading of it, then 7 more permissive documents and 8 more category mismatches once the two deliberately excluded families were admitted.

The third round added a second failure mode worth naming separately: a ledger can carry the right controls and still take its category from the wrong place. _check_column raises one message for five distinct mechanisms, and the ledger inherited one message as one category — so the classification was correct about accept/reject and wrong about why, for a year, in a file whose entire purpose is to be right about why.

cp5 added a third: a comparison surface that gains a member can lose controls for the members it subsumes. Putting message into the BR-V7 dedup key made several key members unobservable at the output, and the cp4 campaign — re-run, not trusted — turned those mutations from caught to survived. The fix was controls that drive the production dedup directly, and the rule is that every earlier campaign is re-run against the new tree, which is what the generated evidence pipeline exists to make cheap.

Final acceptance added a fourth, about the evidence pipeline itself: a campaign's expected catchers can rot between runs without the gate noticing. --validate re-anchored every mutation against the current tree, but it did not ask whether the tests a mutation names still exist or still fail; shadow-cp4's M61 carried two that had stopped failing on main after 4b's promotion. Closed in PR #342: --validate now resolves every named catcher, and the re-run rule is written down.

#260's acceptance work added a fifth, about the comparison itself: a green gate over an empty set is worse than a red one, because a red gate at least says it is awake. The compare driver therefore says out loud how many documents it compared and how many agreed, and the C# samples gate was read from its log rather than from its tick.

#260's sweep added a sixth, about provenance rather than comparison: the environment that records evidence can shape it. A checkout's line-ending setting changes the byte digest of a definition without changing a character; a machine's git identity rides into every commit made from it; a console codepage rewrites a commit message on the way in. None of it is visible to a reader of text on a screen, and all of it is caught only by comparing bytes and commit metadata. The rule: recorded artifacts are produced and compared as bytes; work committed from a local machine enters the tree under a repository identity; a pre-push report scans metadata as well as content; and the fix for the line-ending half is a researched .gitattributes, not an operator's global configuration. PR #344 added the CI half of the same lesson: a workflow's manifest generator wrote the runner's temp path into every document's source — a constant dressed as provenance — and it was caught by the same check on the record before anything was committed; the generator now names a document by its id.

The distinction matters: a ledger cannot report a normative blocker, a preferred sequencing or a remaining acceptance. It carries no project state at all. It is evidence for a claim the status surfaces make, so it is reviewed as evidence — is it derived from the reference or from the port, can it express absence, does every category have a control, and is any family excluded — and when it turns out to be incomplete, the correction is a further census with every result on the record, not a fix to the port.

For how to keep parity work honest — not just its status — see P-022 § Parity-work discipline.

Global PR acceptance packet

Every substantive child PR must state:

Scope:
Explicit non-goals:
Python source of truth:
Frozen fixture:
Fixture regeneration command:
Steady-state test command:
Production dependency changes:
Behavior changes:
Acceptance changes:
Local commands:
GitHub Actions links:
Known deferred cases:

Migration PRs additionally report:

Python-only verdict count:
Rust-only verdict count:
Changed verdict count:
Ordering-only difference count:
Unexplained difference count:

The final value must be 0.

Stop conditions

Stop and redesign the child scope if any of these occur:

  1. A fixture or acceptance must be weakened to make parity green.
  2. Rust intentionally differs without a separate Python-first change.
  3. own-bridge reimplements ownership/lifetime/effect/DI algorithms.
  4. own-codegen starts consuming analysis verdicts.
  5. Release smoke tests use project references/rebuilds instead of the packed .nupkg.
  6. Python and Rust receive different OwnIR bytes in compare mode.
  7. SARIF normalization deletes semantic fields merely to remove a diff.
  8. Performance work begins without a baseline/profile.
  9. A new abstraction layer has no vertical consumer.
  10. Schema, semantics, output and packaging are mixed into one supposedly small PR.
  11. A parity result is reported as complete while known divergence families sit outside the measured set. Excluding them is legitimate; calling the remainder "parity" is not.
  12. A parity category is taken from where the reference raises its error rather than from the mechanism the document violated. One diagnostic covering several mechanisms is normal in a reference written for humans; inheriting it as one category makes the taxonomy decorative.

Note on (1) versus a corrected acceptance: striking a requirement because the tree proves it describes something that does not exist is not weakening it. The distinction is evidence — #256's strikes each carry a measurement, and the surfaces that stated the old acceptance were all corrected in the same move. Narrowing what the reference accepts as a Python-first contract decision with the ledger re-measured against it (#326, then #341) is not weakening either: the reference moved first, and the census followed.

Milestone completion

This roadmap reaches its next major milestone when all are true:

Owen Alpha

  • installable from nuget.org on a clean machine;
  • owen check works on Windows and Linux;
  • immutable GitHub Action tag works from an external repository;
  • SARIF is accepted by GitHub Code Scanning.

Rust vertical

IDE foundation

  • measured latency/memory budgets exist;
  • .own LSP proves cancellation and stale-result correctness;
  • the C# hybrid host has an approved protocol/ADR and a working full-snapshot vertical before delta optimization.

Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions