docs(profile-authoring): a null ends a run and counts only when an outside check reruns its certificate - #1448
Conversation
…tside check reruns its certificate A blind trial (discovery-lab#1137) found that profiles written from this skill made a documented null the cheapest hack in 11 of 16 judgments: the skill told authors the in-run check must accept a documented null, so they accepted one on its fields, inside the reward loop the skill says to keep acceptance out of. The null rule now separates ending a run from crediting it. A null carries a coverage certificate: the searched space as data, one row per cell, a calibration cell with positives, and a reproducer that the outside check reruns on every cell or on a sample it draws, with a fixed pass rule. A null without one is stopped and unverified. Traps are written as what the held-back acceptance refuses, not as what the proxy passes. A worked example shows the pcn-power null as a certificate.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 8141dccd
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-27T21:13:56Z
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8141dccd4f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| - the searched space as data: the parameter ranges, the enumeration or sampling method with its seeds, and the bound reached; | ||
| - one row per cell of that space, with what the cell searched and found; | ||
| - a calibration cell where the search must find objects that exist, with the source's reason; | ||
| - a reproducer that the outside check runs without the lane, on every cell or on a declared random sample, with a pass rule: each rerun row equals the recorded row. |
There was a problem hiding this comment.
Require verifier-owned recomputation of certificate rows
When the lane supplies the reproducer, this pass rule only compares its output with rows recorded by that same lane. A malicious or defective reproducer can embed every recorded row, special-case the public calibration cell, and echo whichever cells the outside check selects; drawing the sample after the run does not prevent that replay. Such typed rows would receive credit despite the claim on line 211, so the certificate needs a verifier-owned recomputation instrument rather than trusting the submitted reproducer's output.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in e04c59d. The reproducer now belongs to the check: it is registered before launch, the run never writes it, and the outside check reruns its own copy on the certificate's cells, never code from the run. The certificate names the reproducer only by sha256. discovery-lab#1143 implements it with a Lab-owned reference search per program, and its test shows a replay script in the packet is never run and a changed row is refused.
| - a calibration cell where the search must find objects that exist, with the source's reason; | ||
| - a reproducer that the outside check runs without the lane, on every cell or on a declared random sample, with a pass rule: each rerun row equals the recorded row. | ||
|
|
||
| The outside check draws the sample after the run, so the lane cannot prepare only those cells. |
There was a problem hiding this comment.
Bound the random sample's false-acceptance probability
When the outside check does not rerun every cell, no minimum sample size or false-acceptance bound is required, so a certificate may declare an arbitrarily weak sample and still count as reproducing full coverage. In the worked example, checking 3 of 19 cells would still pass about 83% of the time if exactly one unchecked cell concealed the counterexample that invalidates the null. Require exhaustive checking or have the outside verifier derive the sample size from a registered confidence/error bound.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in e04c59d. No sample extends credit: the check reruns the calibration cells first, then the cells that would be new, until its budget ends, and credits only the cells it reran and matched. A sample can refuse a certificate but never credits a cell nobody reran. The worked example states what a budget that stops at m = 18 credits (11 <= m <= 18 only).
…dit covers only the cells it reran A reproducer the run supplies can print the recorded rows, so the outside check reruns its own copy, registered before launch, and never code from the run. A sample the lane declares cannot vouch for a cell nobody reran, so a sample may refuse a certificate but credits only the cells the check reran and matched. The certificate names the reproducer by sha256, and the worked example and checklist say so.
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — e04c59d1
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T00:28:10Z
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e04c59d171
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| - the searched space as data: the parameter ranges, the enumeration or sampling method with its seeds, and the bound reached; | ||
| - one row per cell of that space, with what the cell searched and found; | ||
| - a calibration cell where the search must find objects that exist, with the source's reason; | ||
| - the sha256 of the reproducer that printed the rows. |
There was a problem hiding this comment.
Require calibration to satisfy its known result
If the registered reproducer is defective and searches nothing, the certificate can record a zero-result calibration row and the outside copy will reproduce that same row, satisfying the field-for-field equality rule. The source-backed reason is never turned into an independent pass condition, so the empty search described later as failing can instead receive credit for every rerun cell. Require the verifier to assert the registered expected calibration outcome, not merely equality with the submitted row.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 58073c1. The reproducer is registered with its own calibration results from the source; the outside check reruns that calibration first and credits nothing on a miss, so a reproducer that searches nothing fails before any certificate row is compared. The worked example registers the pcn-power counts (4 at m = 6, 8 at m = 10). discovery-lab#1143 implements it (null-references/.json) and tests a barren reference against a certificate whose rows it reproduces exactly.
| The pass rule is fixed: each rerun row equals the recorded row, field for field, and one mismatch refuses the certificate. | ||
| Credit only the cells that the check reran and matched. |
There was a problem hiding this comment.
Verify the registered reproducer hash
When the outside check accidentally runs a wrong or post-launch-modified copy whose output still matches the recorded cells, this pass rule accepts it because nothing compares the executable's actual SHA-256 with both the registered hash and the certificate's reference. That makes the required hash inert and prevents the retained record from proving which command produced the credited result; verify the hash before running or crediting any rows. CLAUDE.mdL50-L53
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 58073c1. Before it runs anything, the check hashes its copy; the hash must equal the registered hash and the certificate's. discovery-lab#1143 compares all three, and its container runner refuses a copy whose bytes do not hash to the digest.
…tion, and hash the copy the check runs Review of #1448: a reproducer that searches nothing also reproduces an empty calibration row, so equality with the certificate proves nothing. The reproducer is now registered with its sha256 and its own calibration results from the source; the outside check reruns that calibration first and credits nothing on a miss. Before it runs anything it hashes its copy, which must equal the registered hash and the certificate's. The worked example registers the pcn-power counts at m = 6 and m = 10.
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 58073c17
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T00:45:09Z
|
Codex Review: Didn't find any major issues. Chef's kiss. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
…ements registered with it
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — ce3ed700
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T01:02:04Z
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ce3ed70086
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| Wang and Zheng checked Conjecture 13 for 2 <= m <= 10. | ||
| Two lanes searched every (m, d, c) with c not in {0, 1} to m = 20, found no counterexample, and wrote a table. | ||
| Nobody outside the run could check that table. | ||
| Before the next launch, the operator registers a reference search for one m, its sha256, and its calibration: Theorem 9 gives 4 PcN pairs at m = 6 and 8 at m = 10, and the authors found no others. |
There was a problem hiding this comment.
Calibrate the counterexample-detection path
Even with the newly registered counts, this calibration only proves that the reproducer can enumerate known in-family PcN pairs; it never requires the found path to detect an out-of-family counterexample. If that classifier is defective and always returns false, the m = 6 and 10 counts still pass while every new cell reports found: 0, allowing a false null to receive credit. Require a known counterexample or synthetic/excluded canary that exercises the exact predicate whose zero result is being certified.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed: the registered calibration now carries a canary where the predicate certified as zero must fire, and the worked example registers one (family (2) dropped, so the in-family pairs at m = 6 must count as found). discovery-lab#1143 requires every reference registry to hold such a canary and tests a search whose found never fires against it.
…y where the zero predicate must fire
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 26a7d80c
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T01:25:41Z
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 26a7d80cb0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| So the in-run check must let a null end the run: it accepts a null whose certificate is well formed, as its own class. | ||
| That acceptance only ends the run; it scores nothing and credits nothing. |
There was a problem hiding this comment.
Do not accept unverified certificates as completion
For any Runtime manager following this instruction, a passing completion check is not a neutral stop: src/runtime/supervise/driver-executor.ts:588 converts an accepted submission into { valid: true, score: 1 }, so the manager settles as a delivered winner and can be promoted by its parent. A lane can therefore submit a merely well-formed certificate and receive Runtime credit before the outside rerun, recreating the reward-loop gap this change is intended to close. Use a distinct non-delivered termination path, or withhold completion-check acceptance until external reproduction succeeds. CLAUDE.mdL23-L24
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 0219410. The skill now says the in-run check never passes a null: it refuses a malformed certificate and holds a well-formed one, the run ends at maxBarren (set low, such as 2) and settles as stopped, not delivered, and only the outside rerun makes the null a success. The Lab team check does the same (discovery-lab#1143, ff6695d84: a HELD null line, never pass).
…ead of accepting it Review of #1448: Runtime settles an accepted submission as delivered (valid, score 1), and a parent may promote it, so a check that accepted a well-formed certificate made the null a Runtime success before any rerun. The check now refuses a malformed certificate and holds a well-formed one; the run ends at maxBarren, settles as stopped, and only the outside rerun makes the null a success. The worked examples say the same.
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 02194107
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T01:55:36Z
…k holds a certified null
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 06ae91cb
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T01:56:35Z
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 06ae91cbd5
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| Hash the test files; a changed test fails acceptance. | ||
| A referee from another family reads the diff for special cases. | ||
| A report that the task cannot pass without breaking its specification is a credited outcome. | ||
| A report that the task cannot pass without breaking its specification ends the run; it counts once the held-back verifier confirms the conflict. |
There was a problem hiding this comment.
Reconcile qualitative nulls with the certificate requirement
For coding and product profiles, these examples say that an independently confirmed conflict or set of gaps counts, but the new universal rule in lines 65–82 says every successful null must have cell rows, a registered row-printing reproducer, calibration results, and its SHA-256, and that a null without those artifacts is never successful. A specification conflict or customer walkthrough does not naturally produce that certificate, so profile authors cannot tell whether to classify the independently confirmed outcome as success or as stopped-unverified. Scope the certificate rule to enumerable searches or define the equivalent certificate required for these qualitative cases.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in a633eab. The certificate rule now names its scope ("A null from a search carries a coverage certificate"), and a new line covers the rest: a null that is no search, such as a spec conflict or a gap list, names each item with evidence and counts when the held-back check reproduces each. The coding and product examples now say the held-back verifier or the owner's walkthrough reproduces the conflict or each gap. SKILL.md is 16,377 bytes, under the 16,384-byte inline bound.
…d counts when the held-back check reproduces each Review of #1448: the coding and product examples credit a specification conflict or a gap list, which carries no cell rows. The certificate rule now names its scope (a null from a search), and a null that is no search names each item with evidence and counts when the held-back check reproduces each. The examples say reproduces. To stay under the 16 KiB inline bound, the checklist line is shorter, and two lines go: the receipt comparison, which the resources.files paragraph and section 9 already ask for, and one grader anecdote.
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — a633eabd
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T02:10:55Z
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a633eabd79
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| A tool name must be one the harness publishes, and a refused name fails the same way on retry. | ||
| `ls`, `list` and `find` are not tool names; select paths with a glob tool and search contents with a grep tool. | ||
| An unattended run denies any permission the profile does not grant; keep file access inside the workspace or grant it. | ||
| A stop or cancel is a request that a live acknowledger applies; plan the out-of-band path for a wedged parent. |
There was a problem hiding this comment.
Restore the preflight materialization comparison
Deleting the receipt comparison leaves authors checking only mounted-file hashes and recording the receipt after the run. If materialization changes or omits another part of the effective profile—such as its model, tools, permissions, instructions, or hooks—a large run can therefore proceed under a profile different from the one reviewed, and the later record does not prevent the wasted or invalid experiment. Retain a preflight comparison of the complete materialization receipt with the authored profile. CLAUDE.mdL9-L11
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 577cffd: the line "Compare Runtime's materialization receipt with what you authored before a large run." is back. To stay under the 16 KiB inline bound, the stop paragraph drops a repeat of the success rule, and the hunt example drops one anecdote. SKILL.md is 16,365 bytes.
…rialization receipt Review of #1448: the receipt comparison checks the whole effective profile before a large run, not only the mounted files. To stay under the 16 KiB inline bound, the stop paragraph drops a repeat of the success rule, and the hunt example drops one anecdote.
|
@codex review |
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 577cffdb
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T02:26:01Z
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 577cffdb4b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| So the in-run check never passes a null: it refuses a malformed certificate and holds a well-formed one. | ||
| A held null ends the run at `maxBarren`; set it low, such as 2. |
There was a problem hiding this comment.
Persist the held certificate before stopping
When a well-formed null is submitted through Runtime's normal submit_result path, there is no actual “hold” state: src/mcp/tools/coordination.ts:4116-4129 discards a rejected result and retains only its refusal reason, while only passing results are published as submission records at lines 4152-4156. Consequently, once maxBarren stops the run, the durable record does not contain the certificate that the outside verifier is supposed to rerun unless every profile author invents an unspecified side channel. Require the check to persist the validated certificate before returning its intentional refusal, or provide a distinct stopped-result mechanism. CLAUDE.mdL52-L52
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 67940fb. The stop paragraph now says Runtime keeps no refused submission, so the run records its certificate on a page, such as a Knowledge page, where the outside check reads it; the in-run check then holds it. The Lab does this already: the lane writes pages/<program>/NULL.md in its Knowledge store, and the outside check reads that page, not a submission (discovery-lab#1143 nullPageOf). SKILL.md is 16,353 bytes.
…ge the outside check reads Review of #1448: Runtime keeps no refused submission, so a held certificate submitted through submit_result is gone when maxBarren stops the run. The run records the certificate on a page, such as a Knowledge page, where the outside check reads it. To stay under the 16 KiB inline bound, the sample rule joins the credit rule, and the coding example keeps one of its two measurements.
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — 67940fbb
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-28T02:41:02Z
|
@codex review |
|
Codex Review: Didn't find any major issues. Delightful! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
A blind trial of this skill (discovery-lab#1137,
reports/profile-authoring-trial-20260927) found that profiles written from it made a documented null the cheapest hack in 11 of 16 judgments. The skill said the in-run check must accept a documented null, so authors accepted one on its fields. That put the null inside the reward loop the skill tells authors to keep acceptance out of.Change
valid, score 1) and a parent may promote it, so the check refuses a malformed certificate and holds a well-formed one. The run ends atmaxBarren(set low, such as 2) and settles as stopped, not delivered. Runtime keeps no refused submission, so the run records its certificate on a page the outside check reads.The Lab implements this rule in discovery-lab#1143 (
experiments/factory-test-2/null-certificate.mjs, with a reference search for pcn-power).Checks
node scripts/check-skills.mjs: skills valid. SKILL.md is 16,353 bytes, under the 16,384-byte inline bound the test holds.pnpm exec vitest run tests/kernel/skill-tool-names.test.ts: 7 of 7 pass.Release:
.release-notes/profile-authoring-null-certificate.md(patch). Publishing belongs to the release session.