Buy vs. build, part 1: What an app actually sells you

Buy vs. build, part 1: What an app actually sells you

In June, the profession’s two biggest conferences gave you opposite instructions. At AICPA Engage, Nicole Davis and Chris Hervochon’s “Not Another Tech Stack” session argued firms should buy the commodity (general ledger, tax software, payroll) and build nearly everything else, because AI moved the line and what was too expensive to build in 2025 is now a weekend project. A week later, the Scaling New Heights expo floor was 200-plus vendors deep, nearly all selling you something.

Here’s the strange part. Presenters at both events kept conceding the same quiet point: most of the accounting AI apps you’re evaluating are LLM wrappers: thin products built on the same AI models your firm can already rent directly. The vendor sends your data to the same Claude or GPT you could prompt yourself, adds a screen, and charges a subscription.

So the buy crowd and the build crowd share a premise. And if the premise is true, building looks unanswerable. I’ve made that case myself: I replaced a Caseware-format working trial balance with an agent in 10 minutes. (An agent is AI that doesn’t just answer; it carries out a sequence of steps itself.) If the intelligence inside these apps is a commodity, why pay for the app?

Because the intelligence was never the whole product. Unbundle any app subscription and you’re buying five things. The wrapper argument prices exactly one of them.

The AI labs themselves just proved it. In July, OpenAI and Anthropic each began selling a persistent work platform around their models: ChatGPT Work, and Claude Cowork now merged with Claude’s chat app. Interface, workspace, and switching costs included. If the intelligence were the whole product, the people who make the intelligence wouldn’t be building workspaces around it.

The five layers inside every accounting AI app

The intelligence is the layer that reads the document, drafts the email, maps the trial balance. It’s a commodity, and the presenters are right.

The interface is what “adds a screen” undersells. A purpose-built review queue beats a chat box for high-volume work. It’s the difference between a tool your whole team can run and one only your AI Champion can drive. It’s also the most perishable layer: agents are removing the human-drives-every-transaction work it exists for, and AI can now ship a working interface in a weekend. The intelligence commoditized first. The interface is next.

The harness scopes the AI to the right client, keeps client A’s data out of client B’s answers, and controls what the model can see and touch. Point one generic AI at your whole practice and you’ve built the confidentiality problem we covered in June’s access-paradox roundup. A vendor whose harness genuinely solves client scoping is selling something real, something most firms would get wrong on their first build.

The janitor is the layer nobody puts on a pricing page. It decides more buy-vs-build outcomes than the rest combined, so it gets its own section.

The pricing wrapper is the per-client or per-user fee structure itself, which turns out to be a risk contract you didn’t read. That one gets part 2 of this series to itself.

The janitor is the real question

Here’s what the janitor does. Migrates your workflows when the model underneath is retired, a six-to-12-month cycle now. Regression-tests output when a model changes silently. Monitors quality, keeps documentation current as staff change, and owns security. As Davis and Hervochon put it, if you build, no vendor stands behind it.

To their credit, their session named this: the true three-year cost, dev plus maintenance plus time, is on their checklist. But the room heard build. The caveats were the actual question.

Buying doesn’t make the janitor disappear, either. This spring, an AI tool many CAS firms run on shipped degraded output for six weeks before customers, not the vendor, caught it. Vendors get paid to be the janitor; that doesn’t mean they always show up.

But when you build, you are the janitor. Every firm with six abandoned automations knows what happens when nobody owns the mop. That’s not an argument against building. It’s the honest price tag the conference sessions left off the slide.

I own one of these apps, so let’s test mine

uCollect is my product. It collects invoice payments for Xero and QuickBooks practices: direct debit, ACH, and card, triggered by the invoices in the ledger. Run the wrapper test honestly and the intelligence layer is thin: an agent could generate the ABA or SEPA bank file this afternoon.

Here’s what an agent can’t conjure: thousands of stored payment authorizations, direct connections into payment gateways, card data tokenized to PCI standards, and the daily discipline of collecting on the due date and posting the payment back for reconciliation. That’s not intelligence. That’s infrastructure and permission. Apps built on rails survive the wrapper test; apps built on prompts don’t.

One caution even survivors carry: the wrapper test doesn’t stop the platform from absorbing your layer. Xero just built document capture into the platform. Every capture vendor’s roadmap now looks mortal. Vendor mortality is a part 3 problem in this series, but check it before you build a workflow on anyone.

Before your next accounting software renewal

For each app on your list, ask: which layer am I actually renting? If the honest answer is the intelligence, you’re paying subscription prices for a commodity. If it’s the interface, the harness, or the janitor, that fee might be the best money you spend, as long as you know that’s what you’re buying, and how fast its clock is running.

And the fifth layer, the pricing wrapper, is where your vendors are quietly making decisions about your margin. That’s part 2. Until then: which layer is each of your apps actually selling you?

If the janitor question is the one you can’t answer for the apps you already run, start there. The free Vendor Test Pack at theaiaccountant.ai/vendor-test-pack is two prompts and a version-trail template for catching the week an AI vendor quietly gets worse. Run it on one workflow before your next renewal.