
The PCAOB released inspection results for the six largest US firms on 13 August 2026. The Big Four's combined Part I.A deficiency rate was 8%, against 20% a year earlier and 26% the year before that. Deloitte and EY landed at 5%, PwC at 9%, KPMG at 13%.
What That Number Actually Measures
One tier down, BDO came in at 34% and Grant Thornton at 33%.
Worth being precise, because the label invites a misreading. A Part I.A deficiency identifies an instance where the auditor did not obtain sufficient and appropriate audit evidence to support its opinion on the financial statements and internal control over financial reporting.
It is not a finding that the accounts were wrong. It is not a restatement. In most cases the opinion may be perfectly sound. What the inspection found is that the file did not demonstrate it.
That distinction is the entire subject. The regulator is not marking the conclusion. It is marking whether the work behind the conclusion was obtained, documented and assembled well enough for a reviewer who was not there to follow it.
A Fourfold Collapse, and a Tier That Did Not Move
Take the three year run at the top seriously: 26, then 20, then 8. That is not noise, and it is not one good year.
Two honest caveats before drawing conclusions. Inspections are risk weighted rather than random, so the selections are not a neutral sample of each firm's work. And firms differ in client mix, which affects how hard the selected engagements are. Comparing 8% to 34% is not the same as comparing two identical firms.
Even allowing for all of that, a gap of four times is not explained by sampling. Something structural separates the two groups.
Evidence Is an Operating System, Not a Talent
Here is the part worth arguing about. Professional judgment is distributed across this profession far more evenly than the results suggest. There are excellent auditors at firms with a 34% rate and mediocre ones at firms with 5%.
What is not evenly distributed is the machinery around the judgment: whether the methodology forces the evidence to be gathered before the conclusion is written, whether the workpaper shows what was examined rather than what was decided, whether a reviewer two levels up can reconstruct the reasoning without a conversation.
That machinery is what the top tier has spent three years and a great deal of money rebuilding, under considerable public pressure. It is unglamorous work. It is also, on this evidence, the thing that moved the number.
The tier below did not do less auditing. It did less of that.
The Same Failure Mode Below the Inspection Line
Most firms reading this will never see a PCAOB inspector. The mechanism does not care.
Peer review asks a version of the same question. So does a professional liability claim, a fee dispute, a successor accountant, and a client who wants to know why a position was taken four years ago. Every one of those reduces to whether the file can carry the argument once the person who made it is unavailable, busy, or gone.
And the stakes have moved recently. When the first draft of a workpaper, a memo or a research conclusion arrives from a tool rather than from a person, the reasoning behind it was never in anyone's head to begin with. A file that records what was concluded but not what was examined used to be recoverable, because someone remembered. That recovery route closes as more of the first pass gets automated.
Which points somewhere specific. The constraint is rarely that people cannot document their reasoning. It is that assembling the evidence takes long enough that documenting it competes with finishing the job.
Where This Leaves the Evidence Base
MetaWurks works on that half. It ingests the underlying material an engagement runs on, the client's statements, invoices, contracts, correspondence and prior filings, and lets an accountant query all of it in plain English, so establishing what supports a number is a question rather than a morning. Role based access controls govern who can open which client's records, audit logs record who opened what and when, and documents ingested into the platform are not used to train models or exposed to other users.
No software writes the memo explaining why the evidence was sufficient. That remains a professional obligation and a professional judgment. What changes is how much of the working day is spent getting the evidence onto the desk, which is the part that quietly decides whether anyone has time to write the memo at all.
The regulator published two numbers this month, 8% and 34%. The distance between them is not made of talent. It is made of what the file can show.
Join the Conversation
Pick your most complex engagement from last year. If a reviewer opened that file today with no access to the people who did the work, how much of the reasoning could they reconstruct?