Archival reproduction package for the manuscript’s three evaluated benchmarks.
results/extendedeval/is the authoritative 15-model, 246-task ExtendedEval matrix.results/appliedeval/is the authoritative 15-model, 57-task AppliedEval matrix.results/humaneval/task_outcomes.jsonlis the authoritative 15-model, 164-task HumanEval matrix.provenance/humaneval_extendedeval_mapping.csvcontains 164 HumanEval rows: 147 verified pairs and 17 explicit absences. No entry-point inference is used for the absences.- Provisional matrices and completion files are not included. See
quarantine/README.md.
datasets/ contains the exact evaluated JSONL snapshots. provenance/ contains the mapping and schema/checksum metadata. results/ contains task-level outcomes, paired outcomes, aggregate tables, universal-failure evidence, and manuscript reports. scripts/ contains the deterministic validation/reproduction command. figures/ contains derived figure source data and release figures.
Tested with Python 3.12 and the standard library only. From this directory:
python3 scripts/reproduce_all.pyThe command fails loudly on record-count, hash, model, mapping, aggregate, paired-statistic, or saturation-summary mismatches. It writes regenerated tables and validation output under reproduced/.
No API key is required. The reproduction command never calls a model provider.
To generate new model completions, see REPRODUCE_EXPERIMENTS.md and run
scripts/run_experiments.py. This requires a user-owned OpenRouter API key.
New reruns must be stored separately from the authoritative results/ files.