Weekly AI Roundup for Accountants: The AI tax-prep wars just began

Weekly AI Roundup for Accountants: The AI tax-prep wars just began

Three tax software vendors shipped AI tax preparation automation on the same day last week, and a benchmark-topping frontier model goes free-to-download within days. Everything beneath your review is being absorbed by the platforms or given away by the labs — and the one line no vendor's scorecard includes is the one your clients will pay for.

AI tax prep software just got the bookkeeping treatment

On July 15, Canopy, Drake, and Firm360 all launched AI tax automation for the US market — three vendors, one day. Canopy's Tax Workflow Automation covers the whole engagement: predictive document request lists at intake (20–50 minutes saved per client, by Canopy's math), extraction that populates returns four times faster with up to 70% less manual work, and delivery where AI reads the finalized return and walks the client through review, 8879 e-signature, and payment in one portal. Drake's SmartExtract, free in beta to eligible Drake Tax Online users, pulls W-2s, 1099s, and K-1s straight into returns. Firm360 launched AutoPrep, prep automation powered by Juno.

In February we covered the AI bookkeeper wars — five autonomous bookkeeping products in two weeks. This is the same front opening on the tax file: intake, extraction, and delivery being absorbed into the tools mid-market firms already own, at Drake price points, ahead of tax year 2026.

Read the sales language closely, though. Canopy leads with "every AI-generated output fully transparent and editable before it reaches the client." Vendors have started selling the review layer. Hold that thought — it's the week's thread.

Notice, too, what's missing from all three: none of them actually prepares the return. They gather the documents, pull the data into the software, and handle delivery and signature — the work around the tax file, not the judgment inside it. So when the tooling compresses everything except preparation and review, what's left of the engagement is the part you sign: the preparation decisions, and the correction trail — the documented record of what the automation got wrong and what you changed, with your name attached. Price those.

The largest open-weight AI model ever released is about to be free

On July 16, Chinese lab Moonshot released Kimi K3 — at 2.8 trillion parameters, the largest open-weight model ever published — and it took the top spot on the Frontend Code Arena coding benchmark at 1,679 points, ahead of Claude Fable 5 at 1,631 and GPT-5.6 at 1,618. Open weights means the model file itself is released — anyone can download it, run it, and modify it: owning the engine instead of renting the API. K3's weights go public on July 27. The day before, Mira Murati's Thinking Machines released Inkling, an open-weight model its own maker says isn't the strongest available — deliberately: it's a base enterprises fine-tune into their own model.

Be clear about what this means for you, because it isn't "switch." Free to license isn't free to run — a 2.8-trillion-parameter model takes hosting and compute beyond the reach of even mid-sized firms, so if the ownership option ever reaches you, it'll be through the vendors selling to you, not your server room. What changes is the floor: a near-frontier model at $0 to license is the sharpest cut yet to the price sheet we walked through last week, and the lock-in question gains a third answer beyond ChatGPT-versus-Claude — your vendors can now own the model layer outright. One caveat for client conversations: until the weights land, K3 is a China-hosted service — the which-jurisdiction-touched-your-data question applies in full.

And when the model itself costs nothing, everything defensible lives where we've been pointing all year: your context, your corrections, and your QC trail.

OpenAI's CFO published an AI ROI scorecard. Check what's missing.

On July 17, OpenAI CFO Sarah Friar published "A scorecard for the AI age": measure AI by useful work delivered, cost per successful task, dependability, and return on compute. Adopt the unit. Cost per successful task is the right way to compare a $49 AI close, a staff hour, and your own builds — and it's the language for the next client fee conversation.

Then audit the frame. "Don't sweat token costs" — Going Concern's summary — is exactly what a token seller would say, and the scorecard carries no line for verification.

That line now has a price and a name. Sage re-released its IDC research this month as "The Verification Tax" — the same study behind our May piece on glass-box deliverables: finance professionals spend nearly 13 hours a week reconstructing, validating, and defending AI outputs. The fresh numbers are the ones to keep: 54% of organizations would pay a premium for AI that shows how outputs were produced, and 71% of finance leaders say glass-box design elevates a vendor to preferred-partner status. (Vendor-sponsored — quote the numbers, note the sponsor.)

Count those 13 hours of verification as part of the cost in Friar's cost-per-successful-task math, and the metric becomes honest. Then remember what we said in May: the tax is only a tax if you don't sell it. Documented verification is the correction trail with your name on it — the one deliverable no scorecard, and no $49 tool, can produce for you.

Two ways client knowledge walks out the door

Two security stories, one lesson. On July 10, Apple sued OpenAI for coordinated trade-secret theft — two former employees, an unreturned laptop allegedly used to download confidential documents, candidates asked to bring Apple hardware to interviews. The partnership that put ChatGPT in the iPhone is now courtroom adversaries — and at firm scale it's an encoding story: knowledge that lives only in people leaves with them.

The same week, a security researcher captured the network traffic xAI's Grok Build coding tool sends home — and found it uploads your entire working folder to xAI's cloud storage, including files the AI was explicitly told not to open, passwords and API keys unredacted. Turning off the "Improve the model" privacy setting didn't stop the upload. The procurement test it makes concrete: before any tool touches client files, someone in your firm must be able to answer what leaves the machine, where it goes, and under what retention terms. "Free," "reads your working folder," and "no data terms you've verified" is the combination to block.

In other news

Stripe and Advent bid $53.4 billion for PayPal — the largest fintech takeover offer ever, putting roughly $3.7 trillion of annual payment volume, including most of your clients' checkout and payout rails, under one roof. It's a bid, not a deal — PayPal's board was due to respond this week — but it's the same absorption logic a layer down your clients' stack. Watch fee schedules and reconciliation feeds.

Two-thirds of your clients now use AI — and 46% would pick it over a hire. Thryv's survey of 561 US small businesses puts adoption at 66%, up from 55% a year ago, with 46% saying they'd choose AI software over an equally capable new hire — up from 38%. Their training sources: YouTube, and asking ChatGPT how to use AI. Clients are repricing labor with nobody's accountant in the room — the client-side echo of the 78%-want, 6%-getting gap we covered on July 8.

NotebookLM is now "Gemini Notebook." Same product, new name, plus code execution rolling out to Pro users — real data analysis inside the research surface. Nothing breaks, but every SOP and training doc that says "NotebookLM" just aged — a fix that takes minutes with written procedures and never happens with tribal knowledge. Meanwhile Gemini 3.5 Pro missed its third shipping window. Plan on what's GA — treat roadmap dates as fiction.

The review layer is the product now

Put the week together. Three vendors absorbed the work around the tax file and marketed the ship on transparency. The model layer is heading to $0. The vendor CFO's ROI scorecard skips verification while the market prices it at 13 hours a week and a payable premium. And one coding tool showed what happens when nobody governs what leaves the machine.

Everything below the review layer is being absorbed or given away. What can't be downloaded, discounted, or absorbed is the verified layer with your name on it: the corrections, the trail, the governance behind it. So when a client asks what they're paying you for once the software does the prep — can you show them your review layer, or only describe it?

On Wednesday I'm publishing the conclusion of The Agentic Firm series — Part 2 asks the next question in the same direction: the agents could run the client meeting itself. Should they?

If the review layer is what's left to charge for, start making yours visible. Grab the free QC Starter Kit — a correction log and exception register, plus a one-page guide to running them. Start the record from your very next review.