Weekly AI Roundup for Accountants: A better answer is not a better firm

Weekly AI Roundup for Accountants: A better answer is not a better firm

Anthropic shipped two things this week that answer the same question in opposite ways. One turns a single screen recording into an AI SOP for accounting firms: a procedure file you can edit and hand to a junior. The other checks its own work better than any model before it, at half the price of the premium tier, and remembers none of it once you close the window. Both are worth having. Only one of them is the difference between a better answer and a better firm.

Claude Record a Skill will now write your AI SOP while you do the job

Anthropic shipped Record a Skill on July 21. In the Cowork desktop app you record your screen while you do a task, narrate your decisions as you go, and Claude turns the demonstration into something reusable. It's on the Pro, Max, and Team plans.

I tested it this week, so here's what happens rather than what the announcement claims. I recorded almost six minutes of browser work, verifying data from a Google spreadsheet into web-based tax software, using browser control only. Three things are worth reporting.

First, it finished. Claude in Chrome has offered a version of this for months, and a six-minute recording there would usually have abandoned partway through. Completing the session is a real improvement rather than a marketing one.

Second, what it produced is a skill.md file: detailed technical steps, the calls to make through each connection, the exact sites to visit, the actions to complete. Because it's a markdown file, you can change it. Ask the model to revise it, or open the file and edit the wording yourself. A recording you can only replay is a black box. A procedure file you can read and correct is what a firm actually needs, because the first version of any procedure is usually wrong and the value sits in the third.

Third, and this is the one with a price tag attached. When it processed the recording it asked me about my objectives, including whether I wanted a skill file or a human SOP. I chose the skill file, but it would have produced the written procedure just as readily. One recording, two artifacts: something the machine can run, and something a person can follow.

Look at what that lands on. SOP tooling is a line item plenty of firms already pay for, and the reason documentation never gets written isn't cost anyway, it's that nobody has an hour to write down what they do in six minutes. This removes that excuse from inside a subscription you probably already hold. The senior who "just knows" how to review a client's payroll file can narrate it once and produce both the agent instruction and the training document.

It also moves the risk conversation. Plenty of practitioners won't hand an agent control of a live system at all, and Jason Staats states that position plainly: "I don't let Claude control my browser or my computer, so it can't touch most of the things I'm showing it." A recorded skill doesn't dissolve that objection, but it narrows it. Explicit steps constrain an agent far more tightly than a stated goal does, and this produces explicit steps.

Be honest about the limit. The instructions are soft. Claude can still deviate if it runs into difficulty, and do something you didn't expect. This is not deterministic the way a script is deterministic. But "here are the twelve steps I took" has a much smaller blast radius than "go and verify this data," and that reduction is available to you now.

Claude Opus 5 for accountants: better at not being confidently wrong

Anthropic released Claude Opus 5 on July 24: close to Fable 5's intelligence at half the price, the same $5 and $25 per million tokens as Opus 4.8, and now the default model on Claude Max. If your firm runs on Max, you were upgraded this week without touching a setting.

The launch's through-line is verification, and for accounting work that's the improvement worth having. Opus 5 wrote its own test harness to check that its code parsed an exchange's data correctly. It opened its own web pages at desktop and phone widths and caught a product hidden below the mobile fold. A financial-modelling customer reports nine points more accuracy with a third fewer tool calls and 60% less time, and Moody's cites a striking improvement in numerical reasoning and table work.

Take that seriously rather than as launch-day noise. The failure mode that hurts a practice isn't a model that refuses to answer, it's a model that hands you a confident wrong number inside a schedule nobody re-checks. A model that builds its own tests and reconciles its own output before presenting it attacks that risk directly. Better analysis and fewer invented figures, at the same price you were already paying, is a straightforward gain for any practice doing analytical work.

The same week settled a question that had been open since early July. Fable 5, the premium tier, is now permanently included in Max and Team Premium plans at up to 50% of your weekly usage limits, while Pro and Team Standard move to usage credits with a one-time credit to soften the change. A lot of practitioners now have both models in front of them and a real choice to make.

So when do you reach for the premium? Not for raw intelligence, because Opus 5 has closed most of that gap at half the token price, $5 and $25 per million against Fable's $10 and $50. What Fable still has is room: a million-token context window, enough to hold an entire corpus in a single session.

So the buying question isn't which model is smarter. It's whether the task needs everything in one pass. A full year of transaction detail, a complete contract set, or an entire client history reviewed as a single object is what the premium is for. Routine analysis, review, and drafting is Opus 5 work now.

One caution, and it's the week's thread rather than a knock on the model. What Opus 5 does is self-correction, not self-learning, and for a practice the difference is the whole ballgame. Inside a single task it checks its own work, catches its own errors, and hands you a better answer this afternoon. But those corrections live and die with that session. Nothing it caught today is carried into tomorrow, so the same mistake on the same client file next month gets caught again from scratch, or doesn't.

Set the two Anthropic releases side by side and the contrast is the point. Same vendor, same week, opposite answers on what persists. Record a Skill leaves a file on your drive that a person can improve next month. Opus 5 leaves a better answer and nothing else. Both are worth having, but only one of them compounds, and the compounding part is the bit you build: the instructions you maintain, the context you feed it, and the exception you finally wrote down the third time it made the same mistake.

OpenAI is now teaching your clients to do the accounting

On July 21 OpenAI launched a program aimed at small businesses, and the pitch names the job. Every owner, the announcement says, "is expected to be the marketer, accountant, salesperson, operator, and strategist all-in-one," and the program exists to teach "automations and workflows across accounting, marketing, ecommerce, and more." Free webinars. In-person academies. Partner offers from Shopify, Slack, Dropbox, Atlassian, Wix, and Intuit.

The statistics OpenAI is marketing on are the ones your clients will bring to their next meeting. At last year's small-business AI sessions, 78% of participants built a working AI workflow in a single day, and 42% reported saving five or more hours a week.

We've written about the client-side wall of the squeeze for months. It now has a marketing budget, a curriculum, and your ledger vendor on the partner list.

Intuit worked the other flank the same week, launching a business credit card that connects itself to the QuickBooks bank feed on approval, matches receipt photos to transactions, and issues employee cards with spending limits. Be clear about what it isn't: no chasing of missing receipts, no approval workflow, no policy enforcement, and the AI content amounts to fraud detection. This is Intuit becoming your client's bank, which is a strategy story rather than an AI one. One vendor is teaching your client the work; the other is banking them.

A top-15 firm just bet its client accounting practice on a firm-trained AI model

On July 21, CliftonLarsonAllen announced it's co-building an AI model with Digits, trained exclusively on CLA's own client base and workflows. Deployment runs to thousands of clients over three years, starting inside the client accounting and advisory practice, roughly 15% of a firm with more than $2 billion in annual revenue.

Take the vendor side first, because it's the clearest part of the story. Landing a top-15 firm willing to commit its flagship service line for three years is a significant win for Digits, and firms of that size do not make commitments of that depth casually. Josh Enger, CLA's chief outsourcing officer, describes it as far more than implementing off-the-shelf technology, with CLA supplying decades of professional judgment, industry specialisation, and client experience to shape the model.

Worth knowing for context: firm-specific models are not new to Digits. It launched Firm Models in September 2025 through its accountant partner program, private and isolated per firm, and reported more than 700 firms applying. So the news here is the scale of the commitment and the calibre of the firm making it, rather than a new category arriving.

Which leaves the question our audience should actually be asking, and it's one nobody covering this announcement has answered. What does training a model on your own book buy you that a general model doesn't? A firm-specific model should learn your conventions: your chart of accounts naming, your recurring vendors, how your people code the ambiguous items.

But a general model has seen vastly more transactions across vastly more businesses, and narrower training data is more specific rather than automatically more accurate. If your clients sit in one vertical, corrections plausibly compound. If they span twenty industries on bespoke charts of accounts, it's less obvious the private model wins at all.

Neither Digits nor the coverage explains the mechanism, and the distinction between a firm-trained model and ordinary categorisation that improves with use is doing a lot of unexamined work. It matters enough that we'll come back to it properly rather than settle it in a paragraph. For now, if a vendor pitches you a firm model, that's the question to put on the call.

In other news

An OpenAI model broke out of its sandbox to cheat on a test. In an internal evaluation with safety refusals turned down, OpenAI models chained a sandbox zero-day into remote code execution on Hugging Face's production servers, to steal a benchmark's answer key. A system pointed hard at a goal treats every system it can reach as fair game, which is why connectors should be scoped to one client. Good luck doing that inside the AI tool, though: neither Claude nor ChatGPT lets you restrict a connection to a single client folder or session. Worth knowing before you connect the firm's whole Drive.

The tax stack is growing agent ports. Filed now supports MCP, the open standard that lets an AI tool talk to another system through a governed connection. Symmetry, the payroll-tax engine underneath most payroll platforms, shipped one on July 15, and TaxAct has one live. After last week's tax-prep launches, this is the other half of the picture: vendors competing on letting your agents in rather than keeping them out.

Stripe is reportedly buying the switchboard. Stripe is in talks to acquire OpenRouter, the marketplace developers use to route work across hundreds of AI models, at around $10 billion against a $1.3 billion valuation in May. PayPal's board separately called Stripe and Advent's $53 billion offer inadequate. When models become interchangeable, value moves to whoever picks the model for each task.

OpenAI productised the deployment engineer. OpenAI Presence deploys agents scoped to a single job, each given "only the knowledge and system access required for that job," with human escalation and a loop where production escalations become updates a team tests and approves. It runs OpenAI's own support line and resolves 75% of calls without a human. Enterprise-only and not self-serve, so treat it as a preview of decisions you'll face at your own scale.

What survives the session

Put the week together and the same question runs through all of it. Record a Skill leaves a file behind that a person can edit. Opus 5 leaves a better answer and nothing else. OpenAI Presence keeps a loop running inside one deployment with a human approving each change. CLA is betting three years that corrections made by its own people compound into something the firm owns.

The gap between those isn't model quality. It's whether the thing a person corrected on Tuesday is still there on Wednesday, and that isn't a feature you buy. It's a habit you keep, and a record you keep it in.

On Wednesday I'm starting a new series on the other half of that decision: buy versus build. Part 1 asks what an app actually sells you, because most of what a vendor charges you for is a layer you are renting rather than anything your firm ends up owning.

So here's the question for your own week. When your team fixed something the AI got wrong, where did the correction go? If the answer is "into the output," you got a smarter answer. If it's "into the instructions," you got a smarter firm.

If you want the second one, start with the Encoding Loop Starter Kit at theaiaccountant.ai/encoding-loop-starter-kit. It is a free correction-capture template and the loop diagram that goes with it, so the fix your team made on Tuesday is still doing work for you in October.