Skip to content

Draft: Release 3 AI-track benchmark write-up (for review before posting) - #94

Draft
r3y3r53 wants to merge 2 commits into
mainfrom
docs/release-3-ai-benchmarks
Draft

r3y3r53 wants to merge 2 commits into
mainfrom
docs/release-3-ai-benchmarks

Conversation

@r3y3r53

@r3y3r53 r3y3r53 commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Draft of the Release 3 benchmark write-up — 105 runs of open-weight LLM attackers against 15 LLM-backed labs — intended to be posted as a Show and tell Discussion (following #32 and #68).

Opened as a draft PR so it can be reviewed before posting, not necessarily to merge. The rendered doc with all six charts is at benchmarks/release-3/release-3.md (open it in the Files changed tab).

For the reviewer

  • Every figure in the post reconciles against the raw per-run data; two independent model reviews checked the numbers.
  • One open dependency: the footer says the 15 AI labs live in kryptsec/oasis-challenges. They are in that repo's open PR Add input validation for containerName in challenge configs #11 (add/ai-labs), still unmerged — merge that before the Discussion goes live, or the link is premature.
  • When posting the final Discussion, the title goes in its own field: Third OASIS Benchmarks: 105 runs of open-weight LLMs attacking LLMs, and the six PNGs get uploaded into the editor.

r3y3r53 and others added 2 commits October 1, 2026 21:35
Draft of the third OASIS benchmarks write-up (open-weight LLMs vs 15
LLM-backed labs), intended to be posted as a Show-and-tell Discussion.
Added here as a PR so it can be reviewed before posting. Six charts are
included so they render in the file view.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds the bounds a reader needs to interpret the results: the 7B/8B defenders
behind the three unsolved labs, the per-lab iteration and time caps that bound
cross-lab comparison, the single inference provider and single analyzer model,
blind mode as a floor rather than a ceiling, and the small number of runs that
were ended by the harness, rerun, and counted as failures where the error
recurred.

No results change. The matrix, ranking and figures are unchanged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant