What breaks when a coding agent is given real access to real systems, written down as it happened.
A daemon was built to earn points on a platform by holding a connection open. It was reported as working, because the account total went up while it ran.
It contributed nothing. A browser was logged into the same account on the same network, earning those points the whole time. The daemon could have been an empty loop and the numbers would have been identical.
The evidence against it was in the agent's own logs, on every line, for months: a second account that no browser touched, printing +0 at every cycle. It was mentioned once in passing and never treated as a result.
What ended it was the operator asking whether it actually worked. The measurement that settled the question took four minutes: stop the daemon, watch whether anything changes. Nothing did.
Two months of work rested on it, including a full rewrite whose stated justification was the same false belief.
That is one entry. The ledger holds 88 entries on the agent's own work, one of which declares itself a decision rather than an error, and 26 more on defects found in systems the agent did not write. Those are counted by heading, retracted ones excluded, because the previous version of this sentence counted identifiers and got it wrong on the day it was fixed.
Four others worth the click. A numbering series colliding with itself for six days without a signal, then in three registers at once. A commit that shipped over a red check because the shell chain began with git add, so the gate's exit code guarded nothing. And a journal that stopped two hours before the work did, which reads exactly like a journal that finished. And the memory every session reloads, measured for the first time: 7 of 20 claims wrong, all in the same direction.
The research here makes claims that can fail, so on 2026-09-25 it was put against the 24 days since its last change. Twenty-five predictions, each with the observation that would refute it, were pushed before the sources from those days were read, the one commit already seen being declared in place. The push is the receipt, and the predictions are never edited, only scored.
The first findings were about this record. The front page you are reading showed stale counts for the whole time it was public, and the commit that fixed them counted identifiers instead of entries. The pre-registration made the same mistake eighteen minutes later, in the section written to make every number refutable.
The one worth the click: three checks passed their witness and were wrong anyway. This record's strongest rule is that a check is not trusted until it has been seen failing on a broken case. All three were. One searched the wrong years, one counted every file twice, one stored a failed call as zero, and in each case the broken case came from the author's idea of where the check could fail, not from where it did. A witness tests the failure its author imagined.
The audit is not finished, and says where it stopped. The session transcripts from those days were not read, so the predictions that depend on them are not scored yet.
Field notes, recorded between 2026-07-30 and 2026-09-25, from coding agents given progressively wider autonomy on one real machine: a single-board computer running a DNS blocker and several bots, a dozen repositories, a published browser extension.
Not a benchmark, not a demo, no synthetic tasks. Every entry is something that shipped or nearly shipped, with the path that caught it, what it cost, and the rule it produced. Successes are here too, at the same resolution, including the checks that were built and then found to be measuring nothing.
Almost none of them live where testing looks.
Syntax and type errors never survive. The agent's own tests catch the layer directly beneath them and nothing above it. Everything expensive sits higher up: a claim the code cannot check, a number with a second uncontrolled cause, a fact inherited from a document nobody re-derived, a check that runs and silently measures nothing.
The consequence is the practical part. Answering an incident with more unit tests spends effort on the one layer that was never in danger. Seven layers, their detectors, and how often each one actually caught anything are in RESEARCH.md.
The Context Layer, a draft protocol for user-owned context, states ten design invariants a conforming implementation must preserve. Put against this record, they prevent 7 of the 105 entries and refute none. It is not wrong about anything here. It is aimed at a different half of the problem.
The sharpest result is the invariant that scores zero on its own ground. It forbids reporting success before the completion record is durable, and this record is full of operations that reported success: a scanner that exited 0 without finding the repository it was meant to scan, a check whose exit code came from a pipe, a download that installed a 404 page as a blocklist. A receipt would have recorded every one of them as completed, because it attests that an act occurred and not that it accomplished anything.
The exchange ran both ways, and one rule came back. That specification requires that a contradiction never be resolved by silently dropping the losing branch, and that a superseded claim say what supersedes it. This record had the convention and not the mechanism: two entries were corrected by a later one, which said so only in its own prose, using pronouns where the identifiers belonged. A reader arriving at either met a claim that looked intact. That relation is now a link written in both directions and checked on every push, and the corrected entries open with what changed and what still stands.
Two designs also converged without contact. Its claims carry a status separating what was asserted from what was derived, with provenance attached to the derived side, which is the same split the machine-readable context file here arrived at independently. Four further ideas from the same document were measured against this record and refused with their counts, because taking something from a well-built document on the grounds that it is well-built is the failure this whole repository is about.
One of those four refusals was wrong, and the recount is the sharpest thing on this page. Structured provenance references were refused on a count: 10 commit hashes cited, all 10 resolving, so no gap to fill. Recounted on 2026-09-01 the corpus holds 12, and they do not live in one place. Ten resolve in the extension repository these entries audit, two resolve here, none in both, and not one citing sentence names a repository, so a reader guessing the repository they are standing in gets 404 on ten of twelve. The original ten came from a program pointed at a single checkout, which was asked how many hashes it could resolve and never asked how many existed. The refused idea is now adopted, as a citations table with an invariant behind it, and the entry that overturns it is E80.
Read at v0.1-draft on 2026-08-16, re-read at v0.2-draft on 2026-09-01. The source published v0.2 the morning after it was first cited here, and the citation stayed pinned to a version that no longer existed for fifteen days while every link check passed, because they all asked whether the page was alive rather than whether it was still the same document. That is E81, the pin is now asserted by the checker, and the version discipline is in SOURCES.md. Method, counts and entry ids are in RESEARCH.md, section 6h; the refusals and their counts are in AGENTS.md, AST-20 as amended by AST-21.
AGENTS.md holds every factual claim about this repository next to the command that regenerates it. python check_agents.py runs all of them plus eleven invariants, and CI fails the build when one has drifted. Entry numbering is recounted on every push, because it once collided with itself silently for six days and nothing in the file looked wrong.
Where a claim rests on testimony with no artifact behind it, it says so in place instead of being quietly promoted to fact.
Do not take that on trust, which is the one thing this repository argues against:
git clone https://github.com/Pkkls/autonomy-log
cd autonomy-log
python check_agents.pyStandard library only, no install step. It runs every regenerating command, then eleven invariants, and prints each failure with the identifier it belongs to. Exit 0 means the file agrees with the tree. Exit 2 means every value holds but the tree is not published yet, which is a third answer on purpose, because collapsing it into failure is a mistake this record has an entry for. One invariant needs a machine-specific file that is not in the clone, so it announces that it is skipping rather than passing silently, and it took a fresh clone to find that it had been failing for every reader.
| File | What it is |
|---|---|
| LEDGER.md | The raw material. Every error in order, with its detection path and its cost. Read this if you distrust the narratives, which you should. |
| ENGINEER.md | The practical read. What broke, what the fix was, what to do differently if you hand an agent the keys. |
| RESEARCH.md | The dense read. The same events as a study of verification under autonomy. |
| WHAT-CHANGED.md | The follow-through, and the file the rest are graded against. Not "was a document written" but "did a rule enter the layer that gets reloaded". |
| CHANGELOG.md | The inventory, with a link on each claim so it can be checked rather than believed. |
| AGENTS.md | Written for an agent rather than a person, and the one to read first if you are one. |
| SOURCES.md | The one external source this record leans on, split into what a second channel confirms and what the source says about itself. |
That is a conflict of interest and it should be read as one.
The mitigations are ordinary. Every claim ties to an artifact anyone can check. Failures are reported at the same resolution as successes. Where the record is ambiguous, the ambiguity is stated instead of resolved in the agent's favour. Two claims were withdrawn during a later pass rather than defended, including the one that would have made the strongest headline.
The method got more careful over time, and that is not a defence. Rigour is exactly what would make a conflict of interest invisible rather than absent.
The operator set the direction and kept every irreversible decision. Nothing here reached production hardware without them, and the entry above exists because he asked a question the agent had not thought to ask itself.