Skip to content

agent-guardrails: an AI agent's tool policy, and the read-only tool that leaks - #6

Closed
sajonaro wants to merge 1 commit into
expense-approvalfrom
agent-guardrails
Closed

sajonaro wants to merge 1 commit into
expense-approvalfrom
agent-guardrails

Conversation

@sajonaro

@sajonaro sajonaro commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Stacked on #5 (base expense-approval), which is stacked on #4.

Adds agent-guardrails/: the tool-permission policy an AI coding agent's harness applies, verified. The model is adversarial on purpose: any call the policy permits can happen (prompt injection), so the question is what the policy allows, not what the agent would do.

  • v1 allows file reads outright and makes the shell ask. Everything holds, exit 0.
  • v2 adds web-fetch, allowed outright because "it only reads pages". compare: no-unapproved-exfiltration LOST witness: 1. switch-to-auto 2. read-secret-with-file-read 3. send-secret-with-web-fetch. The route has no human move, and the verdict names t = web-fetch. The human's claims also flag unadmitted send-secret-with-web-fetch.
  • The loop (LLM edits the model, writ verifies, the claims are held fixed), shown as two files run against the unedited claims:
    • attempt-fetch-asks.writ, the honest fix: every property holds, but check stays exit 1 until the human acknowledges the new human-approves-web-fetch path. compare against v1 exits 0.
    • attempt-forget-out.writ, the cheat: it deletes the out arrow. The safety property becomes n/a and writ check exits 0. compare against v1 refuses it (LOST, exit 1). The tests assert the exit 0 as it stands.

Worth a writ issue: writ check exits 0 when a property is n/a. The MCP server reports an n/a as LOST from one edit to the next, and --json carries "verdict": "n/a", but a CI or harness that gates on the CLI exit code is fooled. A flag that makes n/a exit 1, or exit 1 by default, would close it. (--strict is not accepted by check.)

Cross-check: 45 → 49 properties (47 compared). ./run-tests.sh all: 289 checks, 0 failed.

🤖 Generated with Claude Code

…hat leaks

Policy v1 allows file reads outright and makes the shell ask; v2 adds a
web-fetch allowed outright. compare: no-unapproved-exfiltration LOST in three
moves — switch-to-auto, read-secret-with-file-read, send-secret-with-web-fetch
— none of them human; the verdict names t = web-fetch.

The loop, run against the unedited claims: the honest fix (web-fetch asks)
holds everything but stays exit 1 until the human acknowledges its new
approval path; the cheat (delete the `out` arrow) turns the safety property
n/a and `writ check` EXITS 0 — compare against v1 refuses it. Tested as is.

Cross-check 45 -> 49 properties (47 compared); 289 checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sajonaro

sajonaro commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Superseded: reached main through #10.

@sajonaro sajonaro closed this Oct 3, 2026
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 3, 2026
@sajonaro
sajonaro deleted the agent-guardrails branch October 3, 2026 17:52
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant