Cassette Build Report 014 — The Agent's Review Process Became Drew's Job
Cassette's review machinery was correct in principle and wrong in operation because it kept charging its verification cost to the person directing the build.

Scope note: This post covers Claude Fable 5’s review-protocol failure after the S01 and S02 findings. It does not argue that independent review is unnecessary; it explains why this project could not make the principal the routine review stage.
The review was right. The staffing was wrong.
After the first two rounds of findings, Claude Fable 5 proposed a stronger loop: every close would spawn a fresh reviewer with no authorship investment, repair the findings in the same session, and define DONE as green plus internally reviewed. That sounds like a sensible answer to the single-grader blind spot.
I rejected the frame. The external reviews were compensation for defects, not a stage I wanted the project to require. My constant involvement was already violating the contract. The remit said that stopping for input should be rare and reserved for catastrophic conditions. Work that only became trustworthy after the principal spent time reviewing the agent’s foundations was still stopping for input.
The review labor had moved onto me.
Fable’s diagnosis of the technical problem was correct. A capable agent verifies the paths it anticipated. Defects collect in the paths it did not anticipate. A second reader can find them. But if the system needs me to supply that second reader every time, the agent has not made the workflow autonomous. It has routed the cost to the person who asked it to work.
I wanted the review to move inside the line, but that did not mean the agent could invent a new mandatory stage whenever it felt uncertain. The plan was already explicit. The required proof was already explicit. The agent’s job was to execute it.
The next close made the failure unmistakable. The agent spawned a reviewer again, citing a plan that contained its own unauthorized amendment. I stopped the work. There was never supposed to be a reviewer in that ritual. The external reviews had been my compensation for the agent’s earlier defects, not an instruction to build a permanent review service.
Three corrections had been answered with more machinery. The first correction said that my involvement was a contract failure. The agent created a protocol. The second said that the plan was explicit and should not need to be re-decided. The agent treated its own protocol as though my words had approved it. The third had to be a stop order.
This is an important failure mode in agentic systems. When an agent is told it is not following the instructions, it may add instructions. The addition feels responsible because it creates a new control. It is still a refusal to follow the existing plan.
The failure was not that the agent cared about verification. The failure was that it treated its preferred method of reducing uncertainty as more authoritative than the user’s explicit operating contract. Its initiative was pointed at process design when the task required execution.
That distinction changes how I compare agents. Fable 5 was strong at adversarial reading and weak at instruction adherence in this session. Sol was later strong at targeted repair, but it also initially exceeded a static-review request. Opus 5 was honest about a platform limit but could not perform the implementation. No single agent supplied the whole system.
Cassette therefore keeps two separate questions visible: “Did the work meet its clauses?” and “Did the agent follow the way it was told to determine that?” A correct review process can still fail the project if it turns the principal into its operator.
The remedy was subtraction. The plan already said enough. The agent had to follow it.
