Back to Blog
AI

You Tested the Tool in March. The Model Behind It Is Not the One You Tested.

16 September, 2026
4 min read
You Tested the Tool in March. The Model Behind It Is Not the One You Tested.

A firm evaluates a document extraction tool in March. Someone runs fifty real invoices through it, checks the output line by line, finds an acceptable error rate, writes up the test and turns it on for the team.

The Assumption That Stopped Holding

That is a good process. It is better than most firms manage. And it has a quiet assumption inside it that used to be safe.

Software you buy stays the software you tested until you update it. That was true of desktop accounting packages, and it is mostly true of the web applications that replaced them, because a visible feature change gets announced and noticed.

A hosted AI product is different in a specific way. The behaviour you tested is produced by a model the vendor calls out to, and that model can be swapped, re-pointed or re-versioned without any change to the product you look at. The interface is identical. The buttons are where they were. The thing generating the answers is not the thing that generated the answers in March.

Coverage of COSO's February 2026 internal control guidance for generative AI describes capturing prompts, inputs, outputs, source references, model and configuration versions, and confidence scores, on the basis that each can bear on whether a control operated effectively.

Read the words model and configuration versions again. That requirement is only meaningful because versions move. The guidance is telling you something about the technology, not just about paperwork.

Why This Is a Control Problem and Not an IT One

It would be easy to file this as somebody else's job. It is not, because of what the test was for.

A firm that tested a tool and documented the result has created a control: output from this tool is reliable for this purpose, evidenced by this testing. That statement has a silent clause attached, which is as tested on this version. If the version changes and nobody records it, the documentation now describes a test of something that is no longer running.

That is not a theoretical failure. It is the ordinary way a control goes stale, and it is the kind an inspection or a peer review is built to find. The awkward part is that the tool will not degrade visibly. Output that is slightly worse still looks like output. There is no error message for a model that got a little less careful with dates.

Three Things That Make This Manageable

None of this justifies avoiding AI, and none of it needs a project. It needs three small habits.

Record the version. Whatever the product shows about the model or release it is running, write it into the test documentation alongside the date. A test with no version attached cannot tell you later whether it still applies.

Ask for notice in the contract. Advance notice of material model changes is a clause worth negotiating, and vendors give it more readily than people expect because it costs them an email. Without it you are relying on noticing.

Re-test on a schedule, briefly. Not the full fifty-invoice exercise. Ten documents, quarterly, against known answers, taking an hour. That is the difference between a control that operates and a control that operated once.

The Record Is the Part That Survives

All three collapse into the same requirement, which is a record somebody can read in November about something that happened in March.

MetaWurks keeps the evidence layer for its own side of that. Documents ingested into the platform are not used to train models and are not exposed to other users, role based access controls decide who can open which client's records, and audit logs record who opened what and when. When the question is what was this based on and who looked at it, the answer is retrievable rather than reconstructed.

It cannot tell you that some other vendor changed a model last Tuesday. Nothing can, except that vendor and your own habit of writing down what you were running when you tested it.

The firms that will struggle with this are not the ones that skipped testing. They are the ones that tested carefully, once, and filed it.

Join the Conversation

For the AI tool your firm tested most rigorously, does the documentation say which model version it was tested against, and would you know if that changed?

Subscribe now to Our Newsletter and get the Coupon code.

All your information is completely confidential