Cassette Build Report 010 — Drew Changes the Baseline
Cassette was never meant to match a datacenter from a laptop; it was meant to expand what the same consumer machine could do with a frontier model on external storage.

Scope note: This post explains the performance comparison I intended for Cassette and why it differs from datacenter parity. It does not report a completed benchmark or claim that the new baseline has already been met.
Claude Fable 5 found a serious problem in Cassette’s acceptance matrix. I did not reject the arithmetic. I explained what the project was actually trying to do.
I was not trying to make a laptop compete with a datacenter. Cassette addresses a different limit: consumer hardware cannot normally run a frontier model because most of it does not fit in memory.
The thesis is to keep most of the model on external storage and let the consumer machine do the work it can do. Kimi K3 names a level of model size and capability. It does not name a fixed revision whose laboratory service I promised to reproduce token for token.
The relevant comparison is the same machine without Cassette. What can the machine run unaided? What breadth and long-tail knowledge become available when a much larger model is present on a physical cartridge? How much of the gap between that native ceiling and the model’s reference experience can Cassette close?
That comparison does not make performance unimportant. It makes the claim honest. First-token time, decode rate, answer quality, and model capacity still matter. The difference is that the value gate measures whether Cassette improves the machine it is intended to serve, while a position gate measures how much of the distance to the reference experience it closes.
Two of Claude’s doubts became artifacts of the old encoding under this clarification. The third became an early experiment with a clear cost. The project did not need to pretend that the consumer path matched a datacenter. It needed to establish whether a drive-resident frontier revision could produce enough value to justify the engineering.
I also made two rulings. The remit would be amended in place rather than annotated around the edges. One-time compilation of a frontier cartridge could use large external compute if that cost was recorded openly. Runtime execution had to remain local.
Those rulings matter because a build can become dishonest in either direction. It can lower the target until a small local model counts as success. It can also preserve an impossible parity requirement that makes every useful experiment fail before it begins. I wanted neither. The target had to remain a frontier-scale model on external storage, and the comparison had to describe the value Cassette could plausibly add.
This was not a change from strictness to permissiveness. It was a change from the wrong comparison to the right one. The user experience still had floors. The model still had to remain consequential. Unfavorable measurements still had to be published.
Claude succeeded by forcing me to see what the written contract implied. I succeeded only when I stopped treating its implication as the thesis itself. That exchange is the kind of work I want from multiple agents: one reads the exact words, another explains the purpose those words were meant to serve, and the build carries both forward.
Cassette’s baseline is now explicit. The project must beat the machine’s unaided ceiling and report how far it remains from the model’s reference. A parity label is not a promise. It is a result the measurements have to earn.
