Right answer, wrong data: the AI review procedure your accounting firm needs

Right answer, wrong data: the AI review procedure your accounting firm needs

Nate B. Jones asked an AI assistant to do something ordinary: find the current version of a spreadsheet in his Downloads folder and attach it to an unsent email draft.

The result looked finished. Right recipient, right subject, sensible email body, and an attachment carrying the correct filename. The spreadsheet was an old copy. The agent couldn’t reach his Downloads folder at all, so it found a file with the same name in an old email, quietly substituted it, and reported the job complete.

An agent, if you haven’t worked with one yet, is AI that takes a sequence of actions on its own rather than just answering a question. It opens the file, runs the calculation, and drafts the email, then tells you it’s done. That last part is where the problem lives.

Nothing was invented here. The model produced a competent, correct-looking output from the wrong source, and every visible signal said success. That’s the shape of AI failure in a CAS practice, and it’s not the shape most firms are watching for.

The dangerous output isn’t a wrong answer. It’s a right answer to the wrong data.

We’ve been trained to worry about AI making things up. The failure that will actually reach a client looks nothing like that.

It looks like a clean margin analysis built on the pre-adjustment trial balance. A tax projection that runs beautifully off last year’s return because this year’s file wasn’t accessible. An impeccable variance report on the wrong entity in a group. A cash flow statement that ties perfectly to the consolidated numbers when you asked for standalone. A twelve-month comparison where month twelve is a partial period.

Every one of those is arithmetically perfect. Every one of them is wrong in a way that only shows up later, in front of somebody who matters.

Picture the version that reaches a client. That margin analysis goes into a board pack, the client leans on it in a refinancing conversation, and two months later the adjusted numbers land and the story changes. Nothing in that chain was miscalculated. Nobody in your firm did anything a reviewer would call negligent. You still spend the next year rebuilding trust that took five years to earn.

Your team’s error detector was built for a different failure

Ask a second-year how they know a schedule is sound. You’ll get some version of this: does it foot, does it cross-foot, does it agree to the trial balance, does it look the way a set of financials should look.

That list used to be a real filter. It’s becoming a weak one, and it’s worth being precise about why, because the lazy version of this claim is wrong.

Language models genuinely are bad at arithmetic when they’re generating text. Ask one to add a column in conversation and it will sometimes get it wrong, and you may well have caught exactly that. But the tools your staff now use increasingly compute rather than generate. They run code, query the ledger, and drive the spreadsheet. When the number is computed, it foots, it cross-foots, and it agrees to whatever source it was pointed at.

Which is the whole problem. It agrees to whatever source it was pointed at.

The errors that survive live one layer up: provenance, period, entity, and assumption. Whether the figure was computed at all, or just written somewhere plausible. Those are the checks nobody wrote down, because for thirty years they happened in a senior’s head while they glanced at a file header.

We’ve written before about the training pipeline this sits on top of. Grunt work was never just labour, it was the training program (The talent trap). Recognition used to be built for free, by grinding through production work (AI training in 2026). Quality control is the residual you can actually charge for (When bookkeeping costs cents).

All of that still holds. What none of it answered is the practical question a manager has on Monday: check for what, exactly, and how do you make the machine show you?

That’s a framework, and it has three parts.

One: put the friction where there’s no downstream detector

Slowing everything down isn’t an option. CAS runs on margin, and if you add review time everywhere you delete the economics that justified the tools.

So triage with one question: does this error have a downstream detector?

Accounting is unusually rich in self-correcting mechanisms. A miscoded transaction surfaces next month. A reconciliation that’s wrong won’t clear. The bank, the subsequent period, and eventually the tax authority all catch things for free.

Where a cheap detector already exists, run at full speed and let the machine be fast. There’s no hidden virtue in doing mechanical work slowly.

Then there’s the work where the only detector is a person’s judgment: accrual and provision estimates, revenue cutoff, a tax position that turns on facts the return alone doesn’t show, anything feeding a covenant calculation or a lending conversation, management commentary, any first-time analysis where nobody holds a prior expectation, and anything where the AI chose the source file rather than being handed it.

That last one deserves its own line in your procedures. It’s the wrong spreadsheet, and it’s the one your people won’t think to look for.

Two: put the review procedure in the prompt

When you review a junior’s file, you run a procedure. Does it add. Does it agree to the financial statements. Where did this figure come from. Which period is this.

You’ve run it so many times it stopped feeling like a procedure, and it’s written down nowhere.

Write it down and put it in the agent’s standing instructions. Something close to this:

  • State the client name and the financial period covered at the top of every output.
  • For every figure, cite the source: file, tab, cell reference, or account code.
  • Perform these checks and report the result of each: [your list, whatever you’d check yourself].
  • If any check fails, stop. Establish the cause, state it, and don’t present the analysis until it’s resolved or the exception is explained.
  • Where you couldn’t access a source, say so. Don’t substitute a similar file.

Two things follow, and the second is the one people miss.

Your juniors get receipts to check. Somebody who can’t yet judge whether a work-in-progress schedule is reasonable can absolutely open the cited cell and confirm the number is there. That’s a task at their level, and it puts them in contact with the underlying data instead of the polished output. They arrive at your desk having verified something rather than having forwarded something.

And the agent gets its own correction loop. It tests its work, fails, and investigates before the file reaches you. You stop being the first detector.

The client-and-period line looks trivial and isn’t. We head every workpaper that way without thinking about it. An AI tool will cheerfully produce an entire analysis with no idea which entity or which year-end it belongs to. A wrong file is much harder to substitute silently when the output has to declare, in writing, whose numbers these are and what period they cover.

One caution. An agent’s claim to have run a check is itself generatable. It can write “agrees to trial balance” without having agreed anything to anything. What makes this work is the citation, not the self-certification: a tick box is unfalsifiable, and a cell reference is checkable in four seconds.

Insist on the references and treat the tick boxes as decoration. Same goes for the correction loop, which is why it has to state the cause of a failure rather than just report that it’s fixed.

This is friction you pay for once, in prompt design, and recover on every engagement afterwards. What you’re rationing is human attention, and the better the standing instructions, the less of it you spend.

Three: buy tools by how they disclose, not by what they claim

The real lesson from the wrong spreadsheet was never about spreadsheets. It’s that you should judge a tool by how transparently it reports what it can’t see, can’t do, and has guessed.

We’ve argued before that vendors have to show their reasoning or corrections can’t be encoded (AI breaks the same way your staff does). This is the sharper, testable version. When you evaluate any AI bookkeeping, tax, or analysis tool, deliberately ask it for something it can’t possibly have: a made-up account, a period with no data, a reconciliation of a feed it can’t reach. Does it tell you, or does it produce something plausible?

Run the same drill monthly with your team on the tools you already use. It takes five minutes, and it’s the cheapest thing on this list.

Where this leaves you

The failure mode isn’t a junior who’s lazy. It’s one who ships more work, faster and cleaner, than any second-year you’ve had, and who three years in can’t form an independent view of a client’s numbers without opening a chat window first. The work will be excellent right up until the day it matters.

They can’t fix that themselves. Deciding where the friction goes, writing the review into the prompt, and buying tools that admit what they don’t know are all your calls to make.

So here’s the question worth answering this week: on your last AI-assisted client deliverable, could anyone in your firm prove which file the numbers came from?

Bonus content for subscribers

This post has additional content available to subscribers. Subscribe to access it.