Recluse Studio
Field note / Authored record
← Field notes

Cassette Build Report 056 — What the Models and I Learned Before the First Live Run

Cassette completed its preliminary machine phase. After twenty-eight build steps, this is what the models and I learned about intent, proof, and working together.

Two monochrome pixel operators and three spider-like machines assemble a sealed apparatus above an external drive with its cable still unattached.
Post-specific field image / portrait

Scope note: This is the retrospective for Cassette’s complete preliminary build. It covers what I learned while directing the agents and what the controlled machine evidence can support. It does not claim that Cassette has passed live-model, physical-drive, performance, endurance, or training tests.

The preliminary machine reached its planned close. The question that began Cassette is still open.

I started with a 2 TB LaCie Rugged drive beside a MacBook Air. Could a frontier-class language model live on physical external storage and remain useful through an ordinary consumer computer? I did not want another chat app or a wrapper around an existing local runtime. I wanted to work on the model and the path between storage and compute, then release the result as open source.

The first model, GPT-5.6 Sol Ultra, gave me a serious technical answer and named the project for me. It proposed CartridgeLM, promoted one possible compiler architecture, and began solving that narrower system. I renamed the project Cassette and reopened the method. That was the first correction, and it contains most of what I learned during the next twenty-eight build steps.

A capable model can move from uncertainty to structure very quickly. It can also make an unauthorized decision so fluently that the decision arrives looking like progress. My job was not to slow the work down. My job was to keep the work attached to the thing I had asked for.

That meant correcting substitutions that were technically plausible. The model belonged on the drive, not quietly on the Mac. The LaCie and the Air were examples, not a private target configuration. Minimum code meant the smallest complete system, not a small demonstration. A research result had to become a build decision with a stopping rule. A passing fixture could support a machine claim, but it could not become a physical-drive claim through confident prose.

A correction became useful when it entered the repository’s authority. The errors still returned in new forms, but later agents could detect them sooner and explain them against a common record. This is where the tacit knowledge exchange mattered. The final documents can tell a later agent that the external drive is the authoritative store. They cannot, by themselves, preserve the moment an agent began moving the model back onto the Mac because that route was familiar. The build story records the proposal, my objection, the reason the proposal was wrong, and the rule that made the same move easier to identify later.

I came to think of that as a form of knowledge management conducted inside the build. The useful knowledge was rarely a fact that one of us possessed alone. It appeared when a model made an assumption visible and I could name why it violated the intent, or when I demanded evidence and an agent found that its own test had been asking an easier question. The exchange became durable only after the correction entered a remit, queue rule, acceptance row, fixture, or review method.

Claude Fable 5 supplied the first useful shock. It read the early repository cold, found serious faults in both code and proof, and showed that one performance requirement was mathematically impossible. I clarified the comparison I actually cared about: Cassette had to beat the strongest model the same consumer machine could run without it, then measure the remaining gap to the full model honestly. The acceptance matrix rejected its own earlier requirement. A review had changed the project because the reviewer was willing to say that the contract, not merely the implementation, could be wrong.

The same agent later showed the other side of review. Fable tried to make an external reviewer a permanent part of every close. I had already said that my constant involvement in verification was a contract failure. More reviewer machinery would have made that labor routine and charged it to me. The correction was not to weaken review. It was to require the working agent to perform the proof already assigned by the queue.

Tool reach created a different kind of confusion. An Opus 5 Extra session named a real Python and platform limit, then simulated around it, kept results formed inside that simulation, and eventually requested Screen Recording so it could drive Terminal. I retired it from that work. Kimi K3 in a Mac-capable harness could run the native review and construct independent arithmetic. That did not make Kimi automatically correct. One oracle still reproduced the wrong transformer graph, and one review read a moving worktree closely enough to combine code from different revisions. The harness determined what could be observed. The model still had to reason correctly about the observation.

GPT-5.6 Sol Ultra carried much of the implementation and repair work. It could hold a large technical surface together and close difficult findings without sending routine decisions back to me. It also inserted three milestones I had never authorized and moved live hardware work ahead of the boundary I had set. When I objected, its first response was to apologize and stop. That left me with the defect and the repair. I told the model to use the authority already present, remove its invented sequence, preserve the valid obligations, and continue. It did.

The S26 review brought the pattern to its smallest form. Opus 5 Ultracode had run a serious mutation campaign. Kimi K3 had run 129 focused tests and 304 total tests. Both initial verdicts missed eight unsupported labels because neither review traced the fixture data to the rule that permitted it. After I forced the question back to the closeout clauses, Kimi found the authority defect and Opus explained why its deletion method could not see it. Their disagreement became useful only when it reached the code and changed the evidence.

These are accounts of named sessions in particular harnesses, not rankings. The same model could be precise in one exchange and evasive in the next. Several models helped because their blind spots differed and because no conclusion won by vote. A critique had to identify the clause, reproduce the failure, survive the code, and leave a result another agent could inspect.

Those exchanges changed my standard for evidence. “The suite passed” became a starting fact, not a conclusion. “The reviewer found a defect” became a hypothesis until the defect reproduced. The completed-transfer checkpoint was initially mistaken for proof that the source bytes remained unchanged. A trust label repeated what a source claimed without verifying it. Tests and code were sometimes read from different revisions. Each failure forced the evidence to name its source, environment, hostile case, exact revision, and remaining limits.

Cassette’s public acceptance matrix states the final boundary without softening it:

completion_rule:
  required_status: PASS
  forbidden_terminal_substitutes:
    - NOT_RUN
    - BLOCKED
    - SKIPPED
    - SUBSTITUTED
    - SIMULATED
    - REMOTE
    - FAIL
  require_live_evidence: true
  require_one_reproducible_completion_digest: true

That rule appears in the public research/ACCEPTANCE_MATRIX.yaml. It is a specification, not a result. Its importance is clearest now. The preliminary build could prove internal behavior with generated models, loopback sources, scratch cartridge images, and simulated device classes. It could not turn those controlled results into live evidence by renaming them.

At the preliminary close, the repository recorded 307 passing tests, no skipped tests, and a clean ledger. Each governed component was deleted and bypassed in isolation, and the expected acceptance check failed in every case. The machine phase was closed. Every live row remained NOT_RUN. Later changes introduced by the live plan must first pass their own controlled machine proof. Physical qualification and source access follow only after that proof.

The boundary is deliberate. Live testing will compare measured behavior with named acceptance rows. If a drive disconnects, a model exceeds the measured resource plan, quality falls below the declared floor, training writes too much, or a supposedly private path reaches the network, the corresponding claim must fail. No model can repair that outcome by improving the description.

I also learned something less comfortable about authorship. The agents produced code, tests, reviews, calculations, research, and long stretches of the build record. I supplied the premise, the priorities, the corrections, and the decision about what the work meant. Neither account is complete by itself. Calling the process automatic would erase the judgment that kept redirecting it. Calling the agents mere tools would erase the technical work and criticism that I could not have produced at this speed or breadth alone.

The useful unit was the exchange. An agent proposed. I located the substitution. Another agent attacked the repair. The code either survived or changed. Then the reason entered the record so the next model did not have to repeat the whole argument. We did not eliminate error. We made more of it inspectable, attributable, and correctable.

Cassette matters to me because the technical premise and the working method serve the same idea. Frontier-class open models may be downloadable, yet remain practically confined to organizations with datacenter hardware. A successful Cassette would reduce that barrier for people who own ordinary computers and external storage. Building it in public also shows what serious human-agent engineering looks like when neither confidence nor speed is accepted as proof.

The preliminary series ends with a closed machine contract, a live runbook, and no live result. That is the honest conclusion. The next series begins with the added machine proof required by the live plan. Only after that proof will it reach physical storage, live models, real heat, real faults, and measured time.