AI transaction categorization reads everything except your instructions

AI transaction categorization reads everything except your instructions

Xero's auto-reconciliation has now coded more than 100 million transactions, and since August 13 it labels every one with the method that coded it: a bank rule, a document match, your client's own history, or a prediction from everyone else's books. We covered the launch, and the two new coding benchmarks it landed beside, in Monday's roundup. That coverage left one question open, and it's the one that matters for AI transaction categorization: practitioners already trust exactly one of those four methods, and it's the least intelligent one.

Ask them. Monica Choo, on Xero's own product forum: "Bank Rules are almost always 100% correct. But the Memory and Matches are inaccurate and need to be redone." The layer built from written instructions is the layer that holds. The layers built from statistical inference, however sophisticated, are the ones practitioners are redoing by hand.

That's not a product complaint. It's a category error, and the whole industry is making it.

AI transaction coding is an intent problem, not a pattern problem

A bank statement line is the thinnest piece of context in accounting. "AMAZON MKTPLACE" gives you a merchant and an amount, nothing more. Books, office supplies, or the boss's groceries? The client knows. Amazon knows. The statement line will never say.

Every coding engine on the market answers that question the same way: by inference. Xero's JAX runs two learning layers, defined in its own words: "Memory – JAX decides based on how you've reconciled similar transactions" and "Prediction – JAX suggests based on how other Xero users reconciled similar transactions."

QuickBooks runs a blend: Intuit's own engineers describe a model built "solely on a company's historical data" combined at the moment of decision with a population model, meaning how everyone else's books code that kind of transaction. And QuickBooks auto-categorization accuracy shows it: bookkeepers are watching the population half win fights it should lose. One told me firsthand about a client whose payments to Delta Dentistry kept coming back coded as travel: the client's own history said dentist, everyone else's books said airline, and everyone else kept winning. Intuit doesn't disclose which layer decides, so the reviewer never even sees the fight.

Digits trains its models on 180 million transactions. Feed in receipts and AP documents and you get richer evidence, but it's still evidence: still a guess about what you meant.

And intent legitimately varies. The same vendor serves different purposes in the same file: this month's Amazon charge is reference books, next month's is a warehouse shelf. It varies even more across entities, which is what makes the everyone-else layer risky: a Home Depot purchase is almost always cost of goods sold for a construction company, and almost always repairs and maintenance for a cafe.

How everyone else coded a vendor is not a substitute for knowing this client and what makes them different. The only way to code these transactions correctly is to know the client's policy, and no amount of chart of accounts automation can infer a policy from patterns. It has to be stated.

Xero AI transaction coding in the field: the evidence runs one direction

A Perth firm, Digit Business, ran a structured ten-week test of Xero's auto-reconciliation across a deliberately mixed client portfolio, with 18 team members logging results. Hit rates came in at 60 to 67% on suitable files, and the failures landed at both ends. New files underperformed because there was no history to infer from: the cold-start problem, which just means the system hasn't seen this client do anything yet. Mature files failed the opposite way, with history applied too confidently to transactions that legitimately vary. One tester put it plainly: "I thought that Xero would not reconcile it since it would not be sure, but it would reconcile based on the previous transaction which would make it incorrect."

Too little learning on one file, too much confidence on the next. Those sound like contradictory complaints, but they're one defect: inference without instruction fails at both ends. Where the instruction layer exists, accuracy holds. That's why the bank rule, the least glamorous object in the entire stack, is the only coding method with a fan base.

The market is conceding the point, one transaction at a time

No vendor says "we read your coding policy." But watch what just shipped. QuickBooks' Accounting Agent now lets you put a question to the client on a specific transaction, and the client's answer feeds back into the category suggestion. Booke AI codes directly from a comment the client leaves on the transaction. Digits accepts firm-level and client-level review policies written in plain English and compiles them into checks that run against every new transaction.

Three vendors, three fragments of the same missing input: declared client context. Digits comes closest, and the mechanics are the tell: your written policy becomes a set of checks that audit the coding after the engine has already decided, flagging exceptions into review. QuickBooks and Booke are reactive, one transaction at a time, after the line has already landed in the file. Not one product gives the coding engine itself a standing statement of how this client's transactions should be coded before it decides: the declared layer everywhere ships as a checker or a correction, never as an input. The industry is reinventing the client context file one transaction at a time, and nobody has noticed that's what they're building.

Nobody is measuring the layer that works

Both of the new benchmarks we covered on August 17 tested models with no client history at all. That's a real limitation, but the sharper point is that even with full history the models would be reading the wrong input class. The benchmark the profession actually needs asks a different question: given a client whose stated convention contradicts both the global stereotype and a lazy reading of their history, does the system honour the instruction? No benchmark tests that, because no coding engine accepts the input it would test.

You can run the missing benchmark yourself, this month, on your own client base. Pick a client with a vendor whose name carries a misleading association but which that client codes idiosyncratically: their Delta is a dentist, their Shell is a consultant, their Amazon is a book supplier. Confirm at least five prior codings, enable Xero's auto-reconciliation on that bank account only, and read the method label when the next payment lands. Memory plus the right code means client history won. Prediction, or a code matching the stereotype, means you've found the failure, and the platform's own UI documented it for you.

Then add two controls. Correct one transaction in-row and watch whether the next instance lands right: that tests the vendor's feedback-loop claim at ground truth instead of taking the release notes' word for it. And filter the Reconciled page by method each month, tracking error rates per method. That converts "the AI keeps getting it wrong" from an impression into a number.

The declared layer is yours to build

The platforms will keep improving inference, and reviewing its output stays your job. But configuration is the half of that job most firms treat as an afterthought. Bank rules are the primitive version of declared context: one line, one account, and you can't even export them from Xero. We've made that argument before: the instruction is the asset, not the output it produces.

The full version is a client context file with a coding layer: a standing, written statement of what this client's transactions mean. We only buy books at Amazon. Transfers between these two entities go to the intercompany loan account. This Delta is a dentist.

Every correction your team makes is a sentence that belongs in that file, and right now those sentences evaporate in chat windows and review notes. The knowledge is in your head and your workpapers; no platform can infer it, and no platform will ask for it until the day one ships the feature and the firms with the file ready win the transition. Building that file, client by client, is exactly what we teach in the Practice Transformation Program at theaiaccountant.ai/transformation.

The vendors have spent two years teaching machines to guess what your clients mean. You already know. Where is it written down?