Back to Blog
AI

An Accounting Agent Made 15,000 API Calls in an Hour. Nobody Had Set a Ceiling.

23 September, 2026
4 min read
An Accounting Agent Made 15,000 API Calls in an Hour. Nobody Had Set a Ceiling.

Mandiant's AI Risk and Resilience report, published in September 2026, contains a short account worth more attention than it has received.

The Category Nobody Budgets For

An accounting agent malfunctioned. It entered a runaway execution loop, made more than 15,000 high-cost API calls in less than an hour, generated approximately $50,000 in cloud charges, and interrupted active business transactions while it did so.

No attacker. No prompt injection. No bad judgment in any interesting sense. A loop.

Firms evaluating AI think about two kinds of failure. The tool gets an answer wrong, or somebody breaks into it. Both are real and both get discussed.

This is a third kind, and it is the one that traditional software trained everybody not to expect. A bug in a spreadsheet macro is bounded by what the macro can reach. A misconfigured report runs once and produces rubbish. The blast radius of ordinary software failure is small because ordinary software does not decide how many times to do something.

An agent does. That is what distinguishes it from automation, and it is sold as the advantage. The same property means a failure does not produce one wrong result. It produces wrong results at the rate the infrastructure will accept payment for, until something external intervenes.

In this case nothing did for an hour.

Two Costs, and the Second Is Worse

The $50,000 is the number that travels, and it is the less serious half.

The report notes the agent interrupted active business transactions. That is a live system, in a working day, with real counterparties, degraded by a loop. For an accounting firm the equivalents are unpleasant to picture: a reconciliation process hammering a client's bank feed until the connection is throttled, a document pipeline consuming an API quota that the payroll run needed at four o'clock.

A firm can absorb an unexpected $50,000 invoice with an argument to the vendor. It cannot as easily absorb a client asking why their systems stopped working because of something the firm switched on.

What a Ceiling Looks Like

The fix is not sophisticated, which is why the omission is embarrassing rather than forgivable.

Every agent gets a hard spending limit set at the billing layer, not inside the agent's own instructions. A limit the agent can reason about is a suggestion. A limit enforced by the account is a limit.

Every agent gets a bounded number of actions per run, and a run that exceeds it stops rather than continues. Mandiant's recommendations point the same way: collect telemetry on agent token use, API calls and access to sensitive assets, so that an abnormal rate is visible while it is happening rather than on the invoice.

And every agent gets its own credentials rather than borrowing a person's, so that the rate of activity can be attributed to it at all.

None of that is an AI project. It is the ordinary discipline of giving something authority, which the profession understands better than most industries and has somehow not applied here.

Why This Argues for Narrow Tools

There is a design conclusion underneath the incident, and it is the reason to be sceptical of products that promise to do everything.

A tool that can only read cannot loop expensively through actions, because it has no actions. A tool that can act has to carry the whole apparatus of limits, monitoring and attribution before it is safe to switch on. Those are different products with different costs of ownership, and vendors have an incentive to blur them.

MetaWurks is the first kind by design. It ingests a client's invoices, contracts, statements and correspondence and lets an accountant query the whole set in plain English. It writes to no ledger, initiates no payment and drives no external system. Documents ingested into the platform are not used to train models and are not exposed to other users, role based access controls decide who can open which client's records, and audit logs record who opened what and when.

That is a narrower promise than an agent that runs your close, and the incident above is a reasonable argument for wanting the narrower promise first.

Fifteen thousand calls in under an hour is not a story about artificial intelligence being dangerous. It is a story about authority without a ceiling, which is an old story with a new participant.

Join the Conversation

For any AI in your firm that can act rather than only answer, what is the hard limit on what it can spend or trigger in an hour, and who enforces it?

Subscribe now to Our Newsletter and get the Coupon code.

All your information is completely confidential