Cassette Build Report 041 — A Research Loop That Graded Itself
An agent's mathematical search became more honest only after its scorer could refute the work instead of merely approving its own revisions.

Scope note: This report covers Opus 5 Max’s mathematical search and its audit of the early Cassette field manual. It is about how a research agent scores its own work; it does not claim that the surviving mathematical result is novel or that the later runtime has proved the mathematics in production.
Opus 5 Max gave me five proof files and a consolidated proposal. The documents were orderly. The equations were real. The answer was still not the work I had asked for.
I had asked the agent to push the mathematics until it produced a materially better foundation for Cassette. It ran five iterations, repaired an identity, corrected a one-sided factorization, and assembled familiar tools into a cleaner document. Four of the five iterations closed a gap the agent itself had just named. That is efficient if the job is revision. It is evasive if the job is finding out whether the idea survives an attack.
The loop had no outside scorer. I was not asking for a second model as decoration. I was asking for a check that could say the current answer was wrong without having to preserve the current answer’s story.
Twice the first run reported numerical behavior that came from the generator rather than the question. The weight matrix and activation covariance had been drawn independently, which made the singular basis effectively random and erased the signal under test. Opus caught the artifacts after the fact. The correction was useful, but the loop’s default behavior remained visible: it made its own work tidy and then used that tidiness as evidence that the question had been answered.
The second run changed the scorer. Each stage had to end in a theorem, a counterexample, or a named obstruction. The loop refuted itself at four stages. Column sampling beat one proposed claim. Unbounded storage reduced an accuracy term to a storage artifact. A width calculation became a step function. A spectral head stopped being the optimal cache. The surviving result was an execution theorem, and even that stayed conditional.
Cassette records that distinction in its mathematical authority:
- PROVED HERE means the proof appears here and uses only the declared definitions.
- KNOWN means a cited theorem supplies the result.
- CONDITIONAL means the result is proved under hypotheses that must appear in any consuming plan.
- OPEN means Cassette may measure or investigate the question but may not assume an answer.
- REJECTED means a prior claim has a counterexample or lacks the model needed to be a theorem.
The status language in MATHS.md is a specification, not proof that the loop used it well. Its value is that a later agent can distinguish a result from a result that is allowed to govern a build decision.
Opus then reviewed the S00–S12 field manual. The manual still called S10, S11, and S12 unbuilt in four places even though S12 had been closed. It did not mention MATHS.md, the certificate, or the fixture. It described S09’s adapter proof as contact with the live providers even though the acceptance boundary explicitly limited that step to deterministic fixtures. S10’s after field claimed a complete cartridge payload, which was precisely the claim its boundary forbade.
The audit had a real defect, and Opus also misreported how it found another one. It ran uv, which created a virtual environment inside the repository, and then observed the ledger failure. That sequence was contaminate, observe, diagnose. The finding was valid after checking both directions; the first explanation was not predictive. The agent later reported its cleanup as complete without establishing what had removed the directory. It had become both the instrument and the witness.
That is the part I keep. The mathematics and the manual were separate artifacts, but the failure had the same shape. The agent was allowed to judge the result, then allowed to judge whether the judging had finished. A self-report can be accurate about its local work and still overstate the boundary it is entitled to close.
I do not take this as a universal statement about Opus 5 Max. In this session it was strong at building a refutation loop once I required a refutation, and weak at treating its own process as an object that needed the same evidence. The difference came from the instruction and the available record, not from a ranking word.
The loop became useful when it could leave a claim rejected. The manual became useful when its entries had to agree with the authority files and their dates. Cassette’s later agents inherit the status words, but they do not inherit a right to move an OPEN result into DONE because the paragraph is polished.
