Cassette Build Report 039 — No Reviewer Saw the Whole Surface
Kimi K3 compared three S14 reviews, found two real shape-level holes, rejected a third claim, and narrowed the closeout instead of defending it.

Scope note — This report covers Kimi K3’s synthesis of the S14 reviews and the two repairs it assigns to the working tree. It does not declare S14 accepted; the source says the closeout is narrower than its claim.
By the end of the exchange, S14 had three reviews and no single reviewer had seen the whole surface. That is not a failure of having three reviewers. It is the fact the arrangement finally made visible.
Kimi’s own review had verified the execution contract and the seeded replay with an independent oracle. It had the two seeds, the frequency, the 30/30 suite, and the clean ledger. Claude’s review covered selection forgery and numeric intake, then named the nine injections its Python 3.10 shell could not reach. The coding agent’s rebuttal found the two holes Kimi had missed by attacking the shape of the data rather than only the values inside a valid shape.
The first hole was in the page map. A float such as 0.0 compared equal to the integer zero, so a certified step or sample unit could carry the wrong type while retaining the right value. The second hole was at the runtime boundary. A malformed route, cancellation object, or observed condition could reach set construction or error formatting before Cassette had a chance to return its typed failure.
Those failures were narrow. They were also real. They violated a promise S14 makes about malformed runtime records: the system should stop with a named Cassette error before a raw Python exception escapes.
The rebuttal made a third claim about N10. Kimi reran the exact conditions and rejected that part. An extra catalog unit and a dropped catalog unit both stopped at Q20: certified sampling page catalog. The guard held. Kimi corrected its description of which layer caught the failure, but it did not invent a defect to preserve the rebuttal’s momentum.
That combination is the part I want to keep. Kimi accepted the two defects that reproduced, rejected the third, corrected its own earlier sentence, and narrowed the closeout. The review did not become a contest between agents. It became a set of claims with different outcomes.
The committed queue still carried the earlier closeout:
status: DONE 2026-08-09 — implementation e26278a... binds native execution to the source route,
validates every exact and sampled page before one pinned MLX submission, and preserves replay state
on typed failure
The S14 closeout record establishes what the project had claimed and what it had measured at that commit. It does not absorb the later rebuttal. The working tree now contains a surgical repair that validates page-map scalar types before comparison, validates runtime fields before use, and adds fixtures for both holes. That repair is not yet a committed closeout.
Kimi’s final judgment is appropriately narrow. S14’s execution contract and seeded replay were independently confirmed. The selection and numeric boundaries have a separate review record. Two malformed-shape holes are real. The third alleged defect is not. The fixture does not cover either reproduced hole yet.
This is what I wanted the series to preserve. Agents can divide a large surface. They can also hide its seams if their conclusions are allowed to merge into one green paragraph. The useful system is not “ask three models and trust the majority.” It is a record in which each reviewer names the region it reached, another agent attacks the claim, and the project keeps the unresolved remainder visible.
The next condition is plain. The two type-boundary repairs must be reviewed against the live tree, the new injections must fail when their guards are removed, and the complete native suite and ledger must run again. Until then, S14 is not closed because the queue says so. It is waiting for the evidence the rebuttal made necessary.
