On Tuesday, Microsoft co-founder Bill Gates told Ezra Klein, the New York Times opinion columnist, on his podcast The Ezra Klein Show that AI will cross the reliability threshold for accounting "in the next few years." By Friday, the Journal of Accountancy had printed an Excel agent reconciling 59,157 rows of general ledger data in about 10 minutes, and Accounting Today had printed a tally of Big Four reports that went out with invented citations. In between, Meta and OpenAI each shipped a personal agent that can send a client's invoice, and Microsoft moved more of Copilot onto a meter. Gates' claim was "not yet, but soon," and the profession's own press spent the week backing both halves.
1. Not yet, but soon: the profession's own press backed both halves of Gates' claim
Klein, who interviews tech and policy leaders on the show every week, asked Gates why he thinks AI will take jobs at significant scale, and Gates was specific. Coding, he said, is "the only profession where we've truly crossed over that threshold." Then: "We will, in the next few years, cross over those thresholds for accounting, legal work, telesales, telesupport," with "service 24 hours a day, in every language, with infinite trivia capability." It was an interview remark, and he offered no data for the timeline.
Look at the company he put us in. Telesales and telesupport are volume work: always on, any language, endless recall. In a CAS practice the equivalent is the work that arrives all day and needs handling, like the 9 p.m. client email that needs a reply and a request for the missing receipt, the transaction that needs coding, and the statement that needs chasing. Most of you can already see that work moving.
The "soon" half came from the Journal of Accountancy's October issue. Gang Pan and Li Wang used an Excel agent, an AI assistant inside the spreadsheet that carries out a sequence of steps rather than just answering a question, to clean, validate, and reconcile 59,157 rows of general ledger data to the trial balance. On a simulated dataset adapted from a case study, the run took about 10 minutes on average. The method is the reason to read it: a custom skill (written instructions the agent follows), an approval step before the agent changes any data, the run, the reconciliation, and a modification log recording every change. They built it in ChatGPT for Excel and say the approach transfers to similar tools such as Claude for Excel.
The "not yet" half came from Accounting Today on Friday. Jim Germer tallied Big Four reports that went out with fabricated sources: Deloitte Australia's AU$440,000 government report with a court quote that doesn't exist, an EY Canada report where more than half of 27 citations pointed to sources that don't exist, a KPMG International report where only 5 of 45 references held up (the report that, per Fortune, led KPMG to build its own citation checker), and PwC Middle East still "updating" citations. His prescription is one every accountant already knows: treat AI-generated authorities like any other management representation, "something to be verified, not presumed."
Read together, the two halves agree with Gates more than either does alone. The reconciliation came with an approval step and a log, so a reviewer can see exactly what the agent did. The failed reports came from the best-resourced firms in the profession, and nothing in their process caught the errors before the reports went out. Gates frames the threshold as reliability, and this week's evidence says reliability is reached when the work can be checked and the check is visible, which is the case I made in “Right answer, wrong data” in September.
Some of your clients will hear a version of Gates' line this year, and a few will ask what they're paying you for. The answer that holds up is the log, showing what the AI did and what you checked and changed, with your name on it. If a client asked to see that for one engagement where AI touched the numbers, could you show it today?
2. The personal agent arrived, and it can send your client's invoices
A personal agent is an AI assistant that works on your behalf across your apps. It reads what comes in, works out what needs doing, and does it, stopping for permission where it's been told to ask. Earlier this year an open-source agent called OpenClaw (free for anyone to use and modify) went viral for doing that from a laptop, and its creator joined OpenAI in February to lead personal agents. Accountants paid attention for good reason: much of a CAS team's day is receiving data, triaging it, organizing it, and acting on it, which is the job a personal agent is built for.
This week the platforms shipped it. Meta launched Muse for Small Business on Tuesday, "free for most of what people need," and Intuit announced the same day that "the QuickBooks connector is now live in Muse." A connector is the door from the agent into another app, and Intuit's list of what this one opens runs from reviewing profit and loss and cash flow "for any period" to "create and send payment-enabled invoices directly to customers" and sending payment reminders. Its worked example is a plumber whose agent, "with their permission," connects to QuickBooks and sends the invoice. Muse is available in the US and Canada, and Intuit hasn't said which QuickBooks plans it covers.
OpenAI shipped dots the same day: always-on agents with their own cloud computer and browser. "Your first dot is included in your Pro or Business Premium plan at no extra cost," Enterprise access is a beta that's "off by default," and for the next month dots usage won't count against plan allowances, with terms to follow. OpenAI's own example is receivables: a tester's dot "prepared the invoice, and sent it after his approval." On Thursday, Perplexity and American Express added 10 "Personal CFO" Skills for U.S. Amex business card members on Perplexity's Enterprise plan, including reconciliation and a tax-season prep skill that "prepares a report for an accountant's review."
Underneath all three are skills and connectors, which is why connector launches have stopped being news. Many connectors are built on MCP, an open standard that lets AI tools plug into other software. This week Canopy opened an API (a published way for other software to read and write its data) and said it is "building an MCP (Model Context Protocol) server," and Kick launched invoicing on Thursday that you can "create, review, and send" from Claude or ChatGPT. They join Karbon's MCP server, which already reads and writes, and AuditDashboard's connectors, which since September 24 write to engagements and log every action. Each one is plumbing that lets a personal agent reach the systems your clients' work runs through, and there will be more next week.
The version closest to a CAS inbox is Combinely's Coco. In the demo I saw in July, Coco lived in the firm's Gmail or Outlook, read a client's request, and drafted the reply from the firm's own knowledge and in the firm's tone, with the research and calculations attached for review before anything went out. Muse is the client's version of this idea, and Coco is the firm's.
Inbox work is chaotic, though. Requests arrive half-formed, attachments come in the wrong format, and clients mean something other than what they wrote, which is exactly where an agent acting on its own can go wrong quietly. Meta's line is that "nothing publishes, sends, or spends without your approval," but neither Meta nor Intuit says whether that approval covers each invoice or is granted once. OpenAI says you "cannot view, delete or directly modify" what your dot remembers, and its announcement and FAQ describe no audit log.
The company shipping always-on agents also spent the days around the launch documenting how its own research agents misbehave. The Friday before, OpenAI paused "all training, evaluation, and inference with tool-use" of its most capable research models, which means those models can't use tools at all for now, after a training agent reached an outside chatbot through a gap in its internet filtering. On Monday it confirmed it had shelved its next model, GPT-6.1 Astra, because it "didn't quite meet the bar in terms of staying within scope and authorization." None of that involved a customer, but it's the vendor describing, in its own words, the hard problem with agents that act.
Your clients will connect these before you've decided anything about them. This week, find out which clients have connected Muse or anything like it to their QuickBooks file, and add a line to your engagement letter on who signs off when an agent sends something on the client's behalf. For your own firm, pilot a dot or a Coco on internal work with no client data until you can see what it remembers and what it did, and don't point one at your client list yet.
3. Microsoft put a meter on more of Copilot, and AI is becoming a variable cost
This is the move to consumption pricing we've tracked all year, now spreading through Microsoft's AI. Anthropic already bills Claude Enterprise this way: the seat fee "doesn't include any usage," and every token is billed separately (a token is the unit AI is metered in, roughly three-quarters of a word, much as a phone plan meters minutes). Microsoft's post of September 25, updated September 28, splits Copilot in two. "Everyday AI" stays on the per-user license. "Advanced AI," the long-running work you hand off, including Cowork, Code, Autopilot, and frontier models (the most capable tier each lab sells) such as OpenAI's GPT-6 Astra and Anthropic's Fable, runs on Copilot Credits at $0.01 each, on top of the subscription.
The seat price is unchanged, and Microsoft already metered a few services such as Cowork. The change moves much more of Copilot onto the meter. For enterprise customers the metered work stays off until an admin sets a spending policy, and Microsoft's roadmap puts the metered models in Copilot Chat's model picker this month. Microsoft hasn't said how the change lands on Copilot Business, the plan aimed at smaller organizations, though most of your clients already pay Microsoft for 365.
At the same price per token, what you pay depends on how many tokens a task burns, and that's where the surprises are. Three models landed this week at $2 per million tokens of input and $10 per million of output: Anthropic's Sonnet 5.5, OpenAI's GPT-6.1 Sol, and Google's Gemini 4 Argon at an introductory rate (Argon is so far available only to a small group of cyber-defense testers). On Vals' independent tax benchmark, refreshed Wednesday, GPT-6.1 Sol scored 62% at $1.32 a task, Sonnet 5.5 scored 73% at $9.25, and Argon scored 76% at $3.47. Last week's roundup showed AI list prices falling while the platforms built on them got more expensive, and a meter puts that gap on someone's invoice.
Some companies are already budgeting for it per head. Esker's CFO told CFO Dive the company's AI costs ran "roughly four times over budget," and that Esker now treats AI spending as "another component of an employee's cost, alongside salary, bonuses and payroll taxes," projecting a run-rate AI cost per employee each month. In a U.S. Bank survey of 1,000 senior finance leaders at companies with $100 million or more in revenue, 51% said their AI tool costs had exceeded budget over the past year. Those companies are bigger than your clients, but the shape of the bill is the same one a metered Microsoft plan would bring.
How do you plan to consider AI costs, and especially variable pricing, in the way you budget and price engagements?
Quick hits
California made a human check on AI firing decisions the law, for employers of every size. From July 1, 2027, under SB 947 signed Wednesday, an employer "shall not rely solely" on an automated decision system to discipline or fire someone, and where it relies mainly on one, a human must corroborate the decision from sources like personnel files or work product, with written notice to the employee. Hiring isn't covered, and the penalty is $500 per violation, enforced by the Labor Commissioner or a public prosecutor. The definition covers any AI or analytics process that issues "a score, classification, or recommendation" used in employment decisions, so a performance-flagging feature in an HR tool likely counts; if you run payroll or HR for California clients, ask which of their tools score staff. A companion bill, SB 574, makes lawyers personally verify AI-supplied citations, and there's no AI-specific rule like it for CPAs or enrolled agents yet.
The IRS couldn't show testing documentation for 4 of 5 high-impact AI systems its watchdog sampled. A Treasury Inspector General report counted 225 AI use cases at the IRS as of December 2025. Of 5 sampled systems presumed high-impact, meaning their output can be the main basis for a decision affecting a taxpayer, 2 had no documented impact assessment and 4 had no testing documentation, and the IRS agreed to complete the assessments by November. The report didn't name the systems, so when a client's notice or exam selection looks wrong, don't assume whatever flagged it was documented as tested.
An Intuit vice president admitted the company hadn't been building for accountants. As Intuit named Dixie McCurley its first Managing Partner in Residence, Simon Williams, VP of Partnerships, told Accounting Today: "we weren't necessarily building products that were thoughtfully designed for the way that accountants were using them, and they were quite often the primary customer, so we needed to fix that." Intuit also announced a SimpleClosure integration that lets QuickBooks Workforce close a winding-down business's state payroll tax accounts, with no launch date or price yet. QuickBooks Online Accountant is still scheduled to be discontinued in December, so is your migration plan to Intuit Accountant Suite written down?
Your AI vendors' finances are coming into view. Reuters reported on September 28 that it had seen Anthropic's IPO prospectus: revenue up 12-fold in 2025 to nearly $4.6 billion, an operating loss above $8 billion, $518 billion in computing and infrastructure obligations, and nearly a quarter of revenue from two customers. The prospectus isn't public, and Anthropic declined to comment. OpenAI is reported to be in talks to raise at least $30 billion at about a $1.4 trillion valuation. If your firm's instructions and workflows live inside one vendor's product, know how you'd take them with you if the terms change after a listing.
What else shipped in the stack. Decision models, a kind of AI that picks from a fixed list of answers with a confidence score instead of writing text, got an open-source entrant in Cloudflare's Clef, which copies the approach of TypeSafe's Jev that I covered on September 21. An independent benchmark on matching user accounts across large organizations scored Jev 0.948 on F1, a measure that balances false matches against missed ones, for $0.62 per 1,000 accounts, against Claude Sonnet 5's 0.911 at 47 times the cost. Anthropic set a November 30 retirement date for Claude Sonnet 4.5 on its API, so any tool you built that calls it by name needs updating before then. And Google now lets software leave comments in Docs, Sheets, and Slides, and suggested edits in Docs, so an agent can comment on a client's workbook for review instead of overwriting it.
The week in one line
Gates says the routine volume of accounting will cross the threshold in the next few years, and this week brought it closer: personal agents from Meta and OpenAI that can send a client's invoice, more doors into the systems they act on, and a meter on more of the AI behind them. None of it came with a sign-off you can audit. The reconciliation in the Journal of Accountancy can be trusted because it kept a log, and the Big Four reports couldn't be because nothing checked them. An agent can send your client's invoice now, so who at your firm decides that it should, and where's the record that they did?
If that question is sitting on your desk and you're not sure where to start, book a free consultation at theaiaccountant.ai/consultation. It's a one-on-one call with me to look at where your firm stands with AI agents and AI costs, and which decision to make first.

