Cassette Build Report 037 — Kimi Checked the Rebuttal Before Writing the Story
Kimi K3 reproduced the coding agent's rebuttal against the live tree, accepted two real holes, rejected a third, and corrected its own account.

Scope note — This report covers Kimi K3’s claim-by-claim response to a rebuttal of its S14 review. It concerns evidence ownership and malformed runtime inputs; it does not declare the unfinished repair complete.
I gave Kimi K3 two documents and one instruction. The first document was the coding agent’s rebuttal of Kimi’s S14 verdict. The second was Claude’s attempted review and the trouble that followed. The instruction was the one Cassette keeps learning to need: a concrete claim is checked against the code, including a claim that criticizes the reviewer.
Kimi did not begin by defending the review. It reproduced the claims against the live tree and wrote the account afterward.
The first claim was about page maps. Kimi’s original battery had tested extra fields, missing fields, skipped steps, atom mismatches, empty and duplicate page lists, and an unsupported sample unit. It had also tested boolean and integer confusion at the selection boundary. It had not placed a float into the page map’s step or sample unit fields.
The difference is small enough to miss and large enough to matter. In Python, 0.0 == 0. The old page-map check compared values before enforcing their types, so step=0.0 and sample unit=0.0 entered the schedule as if they were integers. True and 1.0 were refused elsewhere because a different guard caught them. The rebuttal found a real, narrow hole.
The old comparison was visible in the committed S14 implementation:
if (
row["step"] != expected.step
or row["operation_id"] != expected.operation_id
or row["atom_id"] != expected.atom_id
):
The pre-repair page-map relation is an observed defect in the source state Kimi reviewed. It shows the comparison. It does not show the later type-validation repair, which was still in the working tree when Kimi wrote Entry 39.
The second claim was also true. Several malformed runtime records escaped without a Cassette error because fields were used in sets, dictionaries, and error construction before their types were checked. A None source route raised TypeError: 'NoneType' object is not iterable. An object() cancellation token raised AttributeError. A list supplied as an observed condition raised TypeError: unhashable type.
Kimi had tested well-formed values that were wrong. It had not broken the containers around those values. The rebuttal did. That is the distinction I want a reader to keep: a test can cover a field’s meaning while leaving the field’s shape unexamined.
The third claim did not reproduce. The rebuttal said Kimi’s N10 conclusion was false because an extra catalog unit had been accepted and a dropped unit escaped as ValueError. Kimi ran both conditions. Extra unit 99 was refused with CAPABILITY_MISMATCH at Q20: certified sampling page catalog. A dropped unit was refused at the same guard. No execution occurred. No raw escape occurred.
The rebuttal was still right about one thing. Kimi’s earlier write-up had assigned that protection to the selection boundary. The actual backstop was the construction guard. The guard held. Kimi’s sentence about which layer caught the input was wrong.
That correction is the human part of the technical record. Kimi did not flatten the disagreement into “the rebuttal was mostly right.” It separated the true defect, the false defect, and the inaccurate explanation of a real defense. A review can be correct about an account of a guard while incorrect about the guard’s failure. Both facts belong in the same report.
The coding agent’s rebuttal also exposed the limit of the original review. Kimi had attacked values inside valid shapes. The rebuttal attacked the shape itself. The implementation had invited both kinds of input, but the fixture had only named one.
I do not read this as Kimi becoming infallible. Kimi’s first S14 review missed the two holes and overstated part of its mutation battery. What changed was the method after the rebuttal arrived. The claim had a live reproduction, the reproduction had a named boundary, and the agent kept the result that made its own earlier prose less comfortable.
The next report stays with that discomfort. The missing type checks were real. The mutation summary had a second problem: it counted some guards as defended when the test had not actually exercised them. That is a different failure, and it needs its own account.
