
Go looking for accuracy benchmarks on AI document extraction and the numbers arrive quickly. Most tools now read header fields, meaning vendor name, invoice number, date and total, at 97% or better. Accuracy with a validation layer is put at 95 to 99%. Straight-through processing, the share of documents that need no human at all, is reported between 60 and 75% in some places, 60 to 70% in others, and 85 to 92% by one vendor describing its own production environment.
Four Numbers, Four Different Questions
Every one of those figures was published by a company that sells the software.
That is not an accusation of dishonesty. It is a warning about comparison, because those percentages answer different questions and firms treat them as interchangeable.
Header field accuracy is the easiest metric in the category. Vendor name, invoice number, date and total sit in predictable places on a document that has been through millions of training examples. High accuracy there is real and close to meaningless as a differentiator.
Line item accuracy is the hard one, and it is the one that decides whether a bookkeeper can use the output. A single invoice can carry five to fifty lines with descriptions, quantities, unit prices, tax and discount terms, and the published guidance in this category openly describes line items as the hardest remaining problem.
Accuracy with validation is a different thing again. It describes a system with a human checking step folded into it, which means the number partly measures the checking rather than the reading.
And straight-through processing rate is the only one of the four that is close to a business metric, because it describes how much never reaches a person. It is also the number most sensitive to whose documents were used, which is why the published ranges are so wide.
The Concession Buried in the Marketing
The most useful sentence in the whole category comes from a vendor benchmark page rather than from a critic: most exceptions are not extraction failures, so a better parser on its own will never get you to zero.
That is a significant admission and it points at the actual constraint.
If the tool reads a field perfectly but the invoice references a purchase order nobody can find, that is an exception. If the vendor is not in the system, that is an exception. If the line items are read correctly but the coding depends on knowing what the client did with the equipment, that is an exception. None of those get fixed by a more accurate model, and all of them arrive in the same queue.
So a firm choosing between two tools on the strength of 97% versus 98% is optimising a variable that was not the bottleneck.
The Number Nobody Publishes
What a firm actually needs is two figures, and neither appears on a pricing page.
What proportion of your documents, not the vendor's, reach a person. And how long does a person spend on each one that does.
Multiply those together and you have the real cost of the tool in hours per month. A system at 92% straight-through on documents nothing like yours, with exceptions that take eleven minutes each because resolving them means hunting for a contract, can easily lose to a system at 70% with exceptions that take two.
What to Ask For in a Demo
Three requests, all reasonable and all revealing.
Run it on a sample of your own documents, including the messy client who sends photographs of receipts. Ask for line item accuracy separately from header accuracy. And ask what the exception queue looks like: what lands there, what a person has to do to clear an item, and what information they need in front of them to do it.
That last question is where MetaWurks sits, and it is worth being precise about the boundary. MetaWurks is not an AP extraction engine and does not compete with the tools described above. It ingests the client's documents, the invoices, contracts, statements and correspondence, and makes the whole set queryable in plain English, so that when an exception does reach a person the supporting material is a question away rather than a hunt. Role based access controls decide who can open which client's records, audit logs record who opened what and when, and documents ingested into the platform are not used to train models or exposed to other users.
Extraction accuracy decides how many items reach the queue. What the person can find decides what each of those items costs.
Ask both vendors for the second number. Watch what happens.
Join the Conversation
For whatever document automation your firm uses today, do you know what share of documents reach a human, and how many minutes each of those takes?