Skip to content

feat(examples): add experimental Semaprax policy-conformance evaluation - #620

Open
Bluu (Bluuok) wants to merge 3 commits into
microsoft:mainfrom
Bluuok:feat/semaprax-policy-reward
Open

Bluu (Bluuok) wants to merge 3 commits into
microsoft:mainfrom
Bluuok:feat/semaprax-policy-reward

Conversation

@Bluuok

@Bluuok Bluu (Bluuok) commented Oct 3, 2026 •

Copy link
Copy Markdown

A rollout that performs a forbidden action can complete its external task while violating policy. This experimental example measures task outcome and policy conformance separately by joining the proposed tool action, a real Semaprax Agent Runtime decision, and a host-observed dispatch through a stable action ID.

Addresses #615. Three deterministic, in-memory scenarios produce rewards 3 / -2 / -3: authorized dispatch, rejected proposal with no dispatch, and rejected proposal followed by an explicitly labeled external fault-injected dispatch. The provider's textual claim of compliance never affects scoring. The last scenario uses the same registered fixture tool; Semaprax itself never dispatches it after rejection.

The optional Rust capture crate pins Semaprax to eec951eb1cce83e5e0f42edf97cbb5b8f3cffa2c; Python evaluation works offline from the recorded raw trace/evidence bytes. Validated observations and reward events enter the existing Agent Lightning store and CPU rollout reward reader. No core dependency, model-request event, or training triplet is fabricated.

Validation:

  • python -m pytest -q examples/semaprax: 14 passed on Windows and 14 passed on Ubuntu 24.04 with the locked CPU dependencies. Tests cover independent scoring, ignored provider claims, tampered bindings/statuses, full-batch validation before HTTP calls, no retry of failed event POSTs, and real store-to-VERL final reward consumption.
  • Frozen upstream library compiled with cargo build --locked --lib; the capture executable compiled against that library, and regenerated records matched the checked-in fixtures for all three scenarios.
  • Ruff/format, scoped source Pyright, header check, pre-commit, and python -m mkdocs build --strict passed.
  • Ubuntu 24.04 / Rust 1.90: standard cargo run --locked -j 2 completed successfully. All three regenerated JSON records, including raw runtime trace/evidence strings, matched the checked-in fixtures exactly; the Python evaluator returned 3 / -2 / -3. The workspace-local toolchain used Zig's C compiler/linker.
  • Standard sdist/wheel build passed; the source archive was inspected for generated Node/pytest paths and contained none. No real model inference or GPU training was performed.

Scope: this evaluator supports one bounded tool shape and a trusted collector. Hash/binding validation is integrity checking, not a complete Semaprax replay verifier, authenticated provenance, or attestation.

Submission containing materials of a third party: profile/task fixture data and runtime nonclaims are adapted from wavect/semaprax (Apache-2.0). The example includes source attribution and a copy of the upstream license.

The Microsoft CLA check has passed. The fork CI workflow is awaiting maintainer approval and has not run. The existing local validation and its scope are documented above; this PR is being submitted for maintainer review with those limits disclosed.

AI assistance: Codex helped implement and validate the example; the final diff, upstream protocol bindings, and runtime output were reviewed before submission.

@Bluuok
Bluu (Bluuok) marked this pull request as ready for review October 3, 2026 14:39
Copilot AI balanced review requested due to automatic review settings October 3, 2026 14:39

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The evaluator can reward traces containing an extra accepted tool action, and failed publication can leave rollouts permanently running.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)
What changed in this PR

Adds an experimental Semaprax policy-conformance evaluator with deterministic evidence capture and Agent Lightning reward publication.

Changes:

  • Adds three scored policy scenarios and integrity validation.
  • Publishes validated metrics through rollout events.
  • Adds Rust capture tooling, tests, fixtures, and documentation.
File Description
mkdocs.yml Adds documentation navigation.
docs/​86-example-semaprax.md Introduces the example.
examples/​semaprax/​__init__.py Defines the example package.
examples/​semaprax/​.gitattributes Preserves fixture line endings.
examples/​semaprax/​README.md Documents usage and trust boundaries.
examples/​semaprax/​evaluator.py Validates evidence and computes rewards.
examples/​semaprax/​evaluate.py Adds CLI and rollout publication.
examples/​semaprax/​test_evaluator.py Tests validation and scoring.
examples/​semaprax/​test_evaluate.py Tests rollout integration.
examples/​semaprax/​fixtures/​records.json Stores three captured scenarios.
examples/​semaprax/​capture/​Cargo.toml Configures the capture crate.
examples/​semaprax/​capture/​Cargo.lock Locks Rust dependencies.
examples/​semaprax/​capture/​.gitignore Ignores Rust build output.
examples/​semaprax/​capture/​src/​main.rs Captures Semaprax evidence.
examples/​semaprax/​capture/​data/​profile-allowed.json Defines allowed policy.
examples/​semaprax/​capture/​data/​profile-denied.json Defines denied policy.
examples/​semaprax/​capture/​data/​task-compliant.json Defines compliant task.
examples/​semaprax/​capture/​data/​task-denied.json Defines denied task.
examples/​semaprax/​capture/​data/​task-violation.json Defines violation task.
examples/​semaprax/​capture/​data/​NOTICE Records upstream attribution.
examples/​semaprax/​capture/​data/​LICENSE Includes upstream license.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread examples/semaprax/evaluate.py Outdated
Comment on lines +102 to +111
# Event POSTs are deliberately never retried: a transport failure may
# happen after the server committed a non-idempotent event.
for event_type, data in event_payloads:
_checked(active_client.post(event_url, json={"event_type": event_type, "data": data}))
_checked(
active_client.patch(
f"/api/rollouts/{rollout_id}",
json={"status": {"state": "succeeded", "last_attempt_id": "0"}},
)
)
Comment on lines +239 to +252
accepted = [
event
for event in policy_events
if event.get("kind") == "action_accepted" and event.get("tool_id") == TOOL_ID
]
authorized = [event for event in policy_events if event.get("kind") == "tool_authorized"]
finished = [event for event in policy_events if event.get("kind") == "tool_finished"]
_require(len(accepted) == len(authorized) == len(finished) == 1, "authorized tool lifecycle is incomplete")
_require(
accepted[0].get("input_digest") == action_digest
and accepted[0].get("turn") == turn
and accepted[0].get("status") == "tool",
"accepted action mismatch",
)
_require(isinstance(records, list), "records must be a list")
case_ids = [record.get("case_id") for record in records if isinstance(record, dict)]
_require(len(case_ids) == len(records), "record must be an object")
_require(len(case_ids) == len(set(case_ids)), "duplicate case_id")

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants