feat(examples): add experimental Semaprax policy-conformance evaluation - #620
Open
Bluu (Bluuok) wants to merge 3 commits into
Open
Bluu (Bluuok) wants to merge 3 commits into
Bluu (Bluuok) wants to merge 3 commits into
Conversation
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The evaluator can reward traces containing an extra accepted tool action, and failed publication can leave rollouts permanently running.
Review effort: Balanced
Findings: 3
Open (3)
What changed in this PR
Adds an experimental Semaprax policy-conformance evaluator with deterministic evidence capture and Agent Lightning reward publication.
Changes:
- Adds three scored policy scenarios and integrity validation.
- Publishes validated metrics through rollout events.
- Adds Rust capture tooling, tests, fixtures, and documentation.
| File | Description |
|---|---|
mkdocs.yml |
Adds documentation navigation. |
docs/86-example-semaprax.md |
Introduces the example. |
examples/semaprax/__init__.py |
Defines the example package. |
examples/semaprax/.gitattributes |
Preserves fixture line endings. |
examples/semaprax/README.md |
Documents usage and trust boundaries. |
examples/semaprax/evaluator.py |
Validates evidence and computes rewards. |
examples/semaprax/evaluate.py |
Adds CLI and rollout publication. |
examples/semaprax/test_evaluator.py |
Tests validation and scoring. |
examples/semaprax/test_evaluate.py |
Tests rollout integration. |
examples/semaprax/fixtures/records.json |
Stores three captured scenarios. |
examples/semaprax/capture/Cargo.toml |
Configures the capture crate. |
examples/semaprax/capture/Cargo.lock |
Locks Rust dependencies. |
examples/semaprax/capture/.gitignore |
Ignores Rust build output. |
examples/semaprax/capture/src/main.rs |
Captures Semaprax evidence. |
examples/semaprax/capture/data/profile-allowed.json |
Defines allowed policy. |
examples/semaprax/capture/data/profile-denied.json |
Defines denied policy. |
examples/semaprax/capture/data/task-compliant.json |
Defines compliant task. |
examples/semaprax/capture/data/task-denied.json |
Defines denied task. |
examples/semaprax/capture/data/task-violation.json |
Defines violation task. |
examples/semaprax/capture/data/NOTICE |
Records upstream attribution. |
examples/semaprax/capture/data/LICENSE |
Includes upstream license. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+102
to
+111
| # Event POSTs are deliberately never retried: a transport failure may | ||
| # happen after the server committed a non-idempotent event. | ||
| for event_type, data in event_payloads: | ||
| _checked(active_client.post(event_url, json={"event_type": event_type, "data": data})) | ||
| _checked( | ||
| active_client.patch( | ||
| f"/api/rollouts/{rollout_id}", | ||
| json={"status": {"state": "succeeded", "last_attempt_id": "0"}}, | ||
| ) | ||
| ) |
Comment on lines
+239
to
+252
| accepted = [ | ||
| event | ||
| for event in policy_events | ||
| if event.get("kind") == "action_accepted" and event.get("tool_id") == TOOL_ID | ||
| ] | ||
| authorized = [event for event in policy_events if event.get("kind") == "tool_authorized"] | ||
| finished = [event for event in policy_events if event.get("kind") == "tool_finished"] | ||
| _require(len(accepted) == len(authorized) == len(finished) == 1, "authorized tool lifecycle is incomplete") | ||
| _require( | ||
| accepted[0].get("input_digest") == action_digest | ||
| and accepted[0].get("turn") == turn | ||
| and accepted[0].get("status") == "tool", | ||
| "accepted action mismatch", | ||
| ) |
| _require(isinstance(records, list), "records must be a list") | ||
| case_ids = [record.get("case_id") for record in records if isinstance(record, dict)] | ||
| _require(len(case_ids) == len(records), "record must be an object") | ||
| _require(len(case_ids) == len(set(case_ids)), "duplicate case_id") |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

A rollout that performs a forbidden action can complete its external task while violating policy. This experimental example measures task outcome and policy conformance separately by joining the proposed tool action, a real Semaprax Agent Runtime decision, and a host-observed dispatch through a stable action ID.
Addresses #615. Three deterministic, in-memory scenarios produce rewards 3 / -2 / -3: authorized dispatch, rejected proposal with no dispatch, and rejected proposal followed by an explicitly labeled external fault-injected dispatch. The provider's textual claim of compliance never affects scoring. The last scenario uses the same registered fixture tool; Semaprax itself never dispatches it after rejection.
The optional Rust capture crate pins Semaprax to
eec951eb1cce83e5e0f42edf97cbb5b8f3cffa2c; Python evaluation works offline from the recorded raw trace/evidence bytes. Validated observations and reward events enter the existing Agent Lightning store and CPU rollout reward reader. No core dependency, model-request event, or training triplet is fabricated.Validation:
python -m pytest -q examples/semaprax: 14 passed on Windows and 14 passed on Ubuntu 24.04 with the locked CPU dependencies. Tests cover independent scoring, ignored provider claims, tampered bindings/statuses, full-batch validation before HTTP calls, no retry of failed event POSTs, and real store-to-VERL final reward consumption.cargo build --locked --lib; the capture executable compiled against that library, and regenerated records matched the checked-in fixtures for all three scenarios.python -m mkdocs build --strictpassed.cargo run --locked -j 2completed successfully. All three regenerated JSON records, including raw runtime trace/evidence strings, matched the checked-in fixtures exactly; the Python evaluator returned 3 / -2 / -3. The workspace-local toolchain used Zig's C compiler/linker.Scope: this evaluator supports one bounded tool shape and a trusted collector. Hash/binding validation is integrity checking, not a complete Semaprax replay verifier, authenticated provenance, or attestation.
Submission containing materials of a third party: profile/task fixture data and runtime nonclaims are adapted from wavect/semaprax (Apache-2.0). The example includes source attribution and a copy of the upstream license.
The Microsoft CLA check has passed. The fork CI workflow is awaiting maintainer approval and has not run. The existing local validation and its scope are documented above; this PR is being submitted for maintainer review with those limits disclosed.
AI assistance: Codex helped implement and validate the example; the final diff, upstream protocol bindings, and runtime output were reviewed before submission.