Cassette Build Report 011 — The Ledger Rejects Its Own Requirement
Cassette's acceptance system finally applied its own arithmetic to the parity claim it had been enforcing.

Scope note: This post covers the change from an impossible native-parity row to an explicit experimental gate. It does not claim that F4 or F5 has passed or failed in the completed runtime.
The most important calculation in Cassette was one the ledger had been carrying without applying to itself.
After I changed the baseline, Claude Fable 5 and I revised the remit, the acceptance matrix, the evidence record, the build rules, and the execution queue. The project now had three comparisons: the model’s own reference, the hosted service, and B_native, the strongest model the same consumer machine could run without Cassette.
The release gate also became more exact. A frontier-class release needed absolute usability floors. It needed to show value against B_native. It needed to close more than half the gap toward the model’s own reference if it wanted to claim that position. It also needed a published honesty vector, so hiding an unfavorable measurement would itself fail the release.
Then the ledger applied its arithmetic to its own matrix.
Entry E-011 falsified the native-parity row it had mandated. The calculation showed that an unchanged dense path would touch roughly 139.4 gigabytes per token. At 819 gigabytes per second, that path produced 5.87 tokens per second, below the matrix’s floor of ten.
The response was not to lower the floor or quietly remove the row. The row was reclassified as teacher infrastructure. The compiled frontier revision had to live inside a consumer decode budget of roughly 10 to 15 gigabytes touched per token. F4 and F5 became binding experiments with declared kill criteria.
This is what I wanted the acceptance system to do. A contract that can only confirm its author is not an acceptance system. It is a press release with arithmetic attached.
Claude’s contribution was the pressure to run the numbers against the documents instead of trusting the documents because they contained numbers. My contribution was to distinguish the thesis from the encoding that had made the thesis look impossible. The repair preserved the boundary and changed the test.
That distinction matters to AI engineering. Teams often treat a failed benchmark or an impossible threshold as a reason to choose a smaller model, a shorter task, or a more favorable measurement. Sometimes that is the right decision. It is not the same as testing the original claim. A system should say whether it has changed the target, changed the method, or learned that the target is not physically reachable.
Cassette’s ledger now has to carry those differences. A row can be specified without being implemented. It can be implemented without being measured. It can be measured without meeting the release gate. The public record should not collapse those states into a green label.
The useful result of E-011 was not optimism. It was a cheaper truth. Instead of discovering after years that the native-parity requirement could not be met, the project had an early experiment that could resolve the real bet on small hardware.
I am building Cassette as an open-source system, but openness is not only the license. It is also the willingness to let the project state where its own arithmetic breaks, what that break changes, and what remains worth testing.
The ledger did its job when it rejected its own requirement. That rejection is part of the product’s evidence.
