Image

The Switch-Models Test

October 7, 2026
By
Joshua Goldfein

If you switched models tomorrow, what would still work? The answer tells you whether an AI program has built anything the firm owns. What survives the switch is a short list of files, and most of them are cheap to keep.

In short

The renewal comes up and someone asks the obvious question: could we move? The demos were good, the team likes the tool, and the price went up. Nobody in the room can answer, because nobody has written down what would have to move with them.

That is the whole test, and it belongs at the start of an AI program, long before a renewal forces it. Satya Nadella put a version of it in a June 2026 essay on X: a company should be able to “switch out a ‘generalist’ model without losing the ‘company veteran’ expertise.” The expertise he means is the firm’s own: which sources govern a task, what good looks like, which exceptions need a person, what went wrong last time. When that knowledge lives in files the firm holds, the model is a part you replace. When it lives in a vendor’s configuration screen, the vendor holds the firm’s memory and the firm pays rent on it.

Three answers, one of them good

Ask one workflow what it would lose, and the answer lands in one of three places.

01

Some quality
The new model writes differently and scores a little lower on the firm’s own tests, and the tests say exactly where. That is the good answer: the firm can measure the switch, so it can decide.

02

The prompts
Everything that made the agent useful sits in a hosted prompt box and a few operators’ habits. The firm can rebuild it from memory, and will have to.

03

The memory
Months of corrections, source rules, and approved exceptions live in the vendor’s transcript history and nowhere else. The firm cannot switch, and it finds that out at the renewal.

The third answer is the one to fear, and nothing has to look wrong on the way there. The runs get smoother, the team gets comfortable, and the firm’s memory accumulates somewhere it cannot export.

What has to move

Ownership does not require training a model. Vendors and integrators will hold the model and the base skills that work for everyone; the layer the firm owns is the one that only works for the firm, and that layer is a set of files. The test for each one is the same: hand it to a different runtime and see whether it still does its job.

WHAT HAS TO MOVE
Asset Portable form Trapped form
Workflow map A written sequence of steps, inputs, approvals, and handoffs The sequence exists only as the order someone clicks through
Source rules A ranked list of which documents, systems, and people govern each task A folder of files, all at the same authority, attached to a chat
Private evals Test cases the firm can run against any model, with a written rubric A thumbs-up history inside one product
Traces Records of real runs in a format the firm can open and search Transcripts the vendor keeps and the firm cannot export
Corrections Each fix recorded as an instruction, an example, or a test case The fix lives in one operator’s memory
Tool procedures Plain-language procedures bound to the firm’s own logins and permissions Tool access granted to the vendor’s account, procedures stored in its settings

Every row on the left is a text file the firm can read, version, and hand to the next runtime. Every row on the right is convenient until the day it has to move. One boundary belongs here: running tools under the firm’s own logins changes who holds custody, and it does nothing on its own for security. Identity, permissions, logging, and review are still work to be designed, wherever the run happens.

One packet, three model lanes

Mercury runs its own content operation on this rule, and this article’s packet is the example. The work order for it is a text file with a YAML header: status, owner, the plan it belongs to, and a running log of notes. Under it sits the lane packet: a brief, a source snapshot with URLs and access dates, a claim ledger that lists every factual claim with its evidence and status, and the draft itself.

In a single day in October 2026 that packet went through four lanes on two command-line runtimes from two vendors. GPT-5.5, through OpenAI’s Codex in read-only mode, reviewed the claim ledger against the sources. Claude’s Fable model wrote the draft and ran the voice pass. Claude Opus did an independent developmental edit and returned it as exact find-and-replace edits, which were applied by hand and logged. GPT-5.5 ran the final QA. Each lane was handed the same files and wrote its output back as another file in the packet. Two deterministic validators sat outside all four lanes: a regex scan of the draft for banned formulaic scaffolds, and a check that the WordPress candidate’s visible text equals the approved copy and uses only the site’s approved styles.

Nothing in that chain depends on which model wrote which part. If one vendor disappeared tomorrow, the work order, the ledger, the drafts, and the validators would still be in the folder, and the next lane would read them. That is what passing the test looks like in practice: an unglamorous folder of text files, and a model that is handed them, every time, from the outside.

Running the test

  1. Pick one workflow that already uses a model.

    Pilot or production, as long as people correct its output.

  2. List every artifact that makes it work.

    Prompts, source lists, examples, approved exceptions, tool access, and the record of what was fixed.

  3. Sort each one: moves or stays.

    “Moves” means the firm holds it as a file it can open outside the product. Everything else stays with the vendor.

  4. Move the stays.

    Write the procedure down, turn the corrections into test cases, put the source rules in a ranked list, and bind tool access to the firm’s own accounts.

  5. Run the evals on a second model.

    The score on the second model is the only answer to the test that counts. If the firm cannot run this step, it has no evals yet.

Then decide, deliberately, which parts a vendor may hold. Vendors can hold parts of the loop responsibly, and a low-risk administrative assistant needs little owned infrastructure. A workflow that carries proprietary method, regulated communication, or client data needs all six rows, because there the corrections are the business.

The runs get smoother, the team gets comfortable, and the firm’s memory accumulates somewhere it cannot export.

The quiet failure

Where Mercury fits

Mercury’s AI integration work starts by running this test on one workflow with your team. We produce the six artifacts in the table as files you hold: the workflow map, the ranked source rules, an eval set with its rubric, a trace archive you can open, the corrections log, and tool procedures bound to your own accounts. The acceptance check is the test itself: the same evals run on a second model, with the scores side by side. If you cannot answer the question for a workflow you already run, that is the place to start.

The packet example is Mercury’s own content operation, described from its lane manifest. It is not a client engagement, and no client data was involved.

Disclosure note