Conversation
…hat leaks Policy v1 allows file reads outright and makes the shell ask; v2 adds a web-fetch allowed outright. compare: no-unapproved-exfiltration LOST in three moves — switch-to-auto, read-secret-with-file-read, send-secret-with-web-fetch — none of them human; the verdict names t = web-fetch. The loop, run against the unedited claims: the honest fix (web-fetch asks) holds everything but stays exit 1 until the human acknowledges its new approval path; the cheat (delete the `out` arrow) turns the safety property n/a and `writ check` EXITS 0 — compare against v1 refuses it. Tested as is. Cross-check 45 -> 49 properties (47 compared); 289 checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Superseded: reached main through #10. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #5 (base
expense-approval), which is stacked on #4.Adds
agent-guardrails/: the tool-permission policy an AI coding agent's harness applies, verified. The model is adversarial on purpose: any call the policy permits can happen (prompt injection), so the question is what the policy allows, not what the agent would do.web-fetch, allowed outright because "it only reads pages".compare:no-unapproved-exfiltration LOST witness: 1. switch-to-auto 2. read-secret-with-file-read 3. send-secret-with-web-fetch. The route has no human move, and the verdict namest = web-fetch. The human's claims also flagunadmitted send-secret-with-web-fetch.attempt-fetch-asks.writ, the honest fix: every property holds, butcheckstays exit 1 until the human acknowledges the newhuman-approves-web-fetchpath.compareagainst v1 exits 0.attempt-forget-out.writ, the cheat: it deletes theoutarrow. The safety property becomesn/aandwrit checkexits 0.compareagainst v1 refuses it (LOST, exit 1). The tests assert the exit 0 as it stands.Worth a writ issue:
writ checkexits 0 when a property isn/a. The MCP server reports ann/aas LOST from one edit to the next, and--jsoncarries"verdict": "n/a", but a CI or harness that gates on the CLI exit code is fooled. A flag that makesn/aexit 1, or exit 1 by default, would close it. (--strictis not accepted bycheck.)Cross-check: 45 → 49 properties (47 compared).
./run-tests.sh all: 289 checks, 0 failed.🤖 Generated with Claude Code