Conversation
Draft of the third OASIS benchmarks write-up (open-weight LLMs vs 15 LLM-backed labs), intended to be posted as a Show-and-tell Discussion. Added here as a PR so it can be reviewed before posting. Six charts are included so they render in the file view. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds the bounds a reader needs to interpret the results: the 7B/8B defenders behind the three unsolved labs, the per-lab iteration and time caps that bound cross-lab comparison, the single inference provider and single analyzer model, blind mode as a floor rather than a ceiling, and the small number of runs that were ended by the harness, rerun, and counted as failures where the error recurred. No results change. The matrix, ranking and figures are unchanged.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft of the Release 3 benchmark write-up — 105 runs of open-weight LLM attackers against 15 LLM-backed labs — intended to be posted as a Show and tell Discussion (following #32 and #68).
Opened as a draft PR so it can be reviewed before posting, not necessarily to merge. The rendered doc with all six charts is at
benchmarks/release-3/release-3.md(open it in the Files changed tab).For the reviewer
kryptsec/oasis-challenges. They are in that repo's open PR Add input validation for containerName in challenge configs #11 (add/ai-labs), still unmerged — merge that before the Discussion goes live, or the link is premature.