On Tuesday, September 22, Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output (a token is the unit AI models charge by, roughly three quarters of a word), 20% below Opus 5, and OpenAI released GPT-6 Sol and GPT-6 Luna at half the promotional price of the GPT-5.6 models they replace. The same Tuesday, two tax-software companies published a working Form 1040 calculation engine under an open-source license, and one of the founders told Accounting Today it's "deflationary with respect to service fees and license fees." On Friday, FreshBooks announced a Claude connector that reaches "all active businesses associated with your login," which for an accountant is a different sentence than it is for a business owner. And a survey of 486 firms our size found that 87% of firms using AI have no written policy for it. The price of the parts fell all week; what didn't fall is the subscription on the ledger, which goes up October 1 for Xero subscribers, or the cost of deciding how your firm will use any of this.
1. Two price cuts on one Tuesday, measured with two different rulers
Both vendors cut the per-token price on the same day. Anthropic's Opus 5.5 lists at $4 and $20 per million tokens against Opus 5's $5 and $25, with cached input (text the model has already read once) at 60% less. OpenAI's GPT-6 Sol lists at $2 and $10 where GPT-5.6 Sol had been $4 and $20, and Luna, the small model, at $0.10 and $0.50 where its predecessor was $0.20 and $1.20. Grok 4.7 arrived the day before at $2 and $6. Three vendors, one week, every list price down.
The headline percentages need their denominators. Anthropic's "40% less than Opus 5 on typical workloads" is measured "at default settings," and the default effort setting (how much reasoning the model does before it answers) dropped from high to medium in the same release, so part of the saving is the model doing less. OpenAI's "50%" is measured against the promotional price of GPT-5.6, the model Sol replaces, which had been discounted since August 21. So Anthropic's cut is 20% off a list price, OpenAI's is 50% off a promotion, and neither number is what you pay for a finished piece of work.
An independent benchmark refreshed on Wednesday supplies the number that is, and it's an accuracy number before it's a price. Vals AI's Tax Agent Bench gives AI agents (software that works through a task in steps on its own, searching, reading, and calculating, rather than answering a single question) 391 corporate-tax research questions written by tax professionals, a set of tools, three hours, and a rubric of about 24 checks per question. Claude Fable 5.1 leads at 77.64% for about $13.18 per question, Opus 5 is second at 75.06% for $5.14, and GLM 5.3, a model from the Chinese lab Z.ai whose weights are published for anyone to run, third at 73.09% for $1.80. The new Opus 5.5 came sixth at 70.50%, and at $14.32 per question it was the most expensive model on the board; GPT-6 Sol scored 53.05%. The question the ranking can't answer is whether 77.64% is acceptable at any price, because Vals hasn't scored a tax professional on the same rubric, and until someone does, the number has nothing to be measured against.
Vals is paid by the labs it scores, the runs are single, and cost depends on how much effort each model was configured to use, so treat it as one number pointing the other way rather than a refutation. It's also the only number this week that didn't come from the vendor selling the model.
Two more things from the same benchmark matter more than the ranking. No model clears 80%, and on the stricter measure, every check on a question satisfied, no model reaches 50%. So the best agent passes about 78% of the checks on an average question and every check on about half of them, and on the one public question, twelve of eighteen models scored zero by missing a single must-pass check. A missing required element is the omission failure I described in my September 4 article, "Right Answer, Wrong Data," and it's where review minutes belong.
I've made the cost-per-accepted-result argument in four of the last six roundups, so I'll keep it to one line here: the price of a token is the price of an input, and the price of a correct, reviewed, signed answer is the only number your fee has to cover. The point this week is what happens on the other side of that ledger. Xero's US price rises on October 1 with AI as the stated reason, Intuit's rose on August 1 on the same justification, and the published price of the same intelligence fell by a fifth to a half on Tuesday; what Xero and Intuit actually pay for it is unpublished. If you're willing to route work across models yourself, it got cheaper than that: GLM 5.3 sits two points behind Opus 5 at about a third of the cost per question.
What this means for you. First, put a number on it: pick one recurring engagement and work out what a correct, reviewed, signed answer costs you today, because two vendors just lowered the input cost and you can't tell whether it reached your margin without that figure. Second, at your next platform renewal, ask the question the vendors haven't answered: the published price of the intelligence inside your software fell this week, so what in the price increase reflects that? You may not get an answer, but you'll find out whether the vendor has one.
2. Someone gave away the 1040 engine, and said what it was for
On Tuesday, Filed and Crimson Tree Software released OpenTax, in the project's words "a single binary that runs on macOS, Linux and Windows with no account and no cloud dependency," meaning one program you download and run on your own machine, covering Form 1040 for tax year 2025, checking returns "against the IRS MeF business rule set," and exporting the XML the IRS e-file system accepts. The engine is built as a chain of 186 steps, each feeding the next, so a W-2 box flows to line 1a and on to AGI, with 131 input types covering W-2s, 1099s, schedules, credits, and capital transactions. It's licensed under AGPL v3, an open-source license that requires anyone who modifies the engine and serves it to customers to publish those changes, with a paid commercial license as the alternative. The project says it is "built and maintained by AI agents using official IRS publications as the sole source of truth," with a human reviewing changes before merge.
The founders said what the release was for. Filed's Leroy Kerry: "The core calculation is a solved problem. The incumbents solved it decades ago and have been charging a premium for it ever since." Crimson Tree's Tom O'Sullivan, to Accounting Today: "It's deflationary with respect to service fees and license fees for software… it will drive value in the work that the clients value, not in the work that only the vendors value."
The limits are the project's own, and they're real. Its FAQ asks "Can I file my tax return with OpenTax?" and answers "No." It doesn't transmit returns; the IRS e-file application is "pending," which the project says is "not proof of IRS approval." Validation is 133 benchmark scenarios drawn from IRS training material, with no independent accuracy figure. And it's a US federal 1040 engine, so if you're reading this from Canada, the product doesn't apply to you and the pattern does.
Here's why I think this is the most significant development in AI for tax this year, ahead of anything the platforms have shipped. Tax preparation is the corner of the profession where the calculation engines are proprietary, expensive to rebuild, and until Tuesday unpublished: almost all of the AI investment in tax has gone into the layer around the return, document collection, client portals, and practice management, because nobody outside the incumbents had an engine you could read, run, and check. An open engine changes what's possible for the tools you already use. A document-gathering platform that can already collect and verify a client's source documents can now, in principle, prepare the return, and with its own e-file acceptance, transmit it. Whether OpenTax is good is unestablished; whether an open engine changes what tax software can charge for is the question, and its makers say yes.
Watch the license, because it tells you where the money moves. A vendor that modifies the engine and serves it inside its own product either publishes those changes or buys the commercial license from the two companies that hold the copyright; the toll is on anyone who wants to build on it privately. Filed sells AI tax preparation. So a tax-prep company has opened the calculation layer to make the layer above it, the agent that gathers, reasons, and drafts, the product, and it collects a toll from anyone who follows. The calculation became a commodity on Tuesday, and the company that made it one has already decided where it wants to be paid.
There's a second question in here for the profession, and it's about the screen. OpenTax has no interface at all, and Filed's own product doesn't replace one either: it "populates the return directly inside your tax software," with the preparer reviewing in the screens they already use. Nobody reviews a return they can't see, so the real question is who owns the review surface once the calculation is free, because the forms view is the reviewer's instrument, with prior-year comparison, diagnostics, and the line your eye goes to first. If the engine is a commodity, the interface is the product, and Filed's own site showing its product working inside Drake Tax rather than beside it tells you the vendors know it.
What this means for you. First, it runs on a laptop with no account, so the narrow Monday use is a second-check pass: key a handful of this season's simpler filed 1040s into it, by hand, since it has no interface, and see what disagrees. Where the engine and your software differ, one of them is wrong, and finding out which is worth an afternoon. Second, the bigger question, and one to start on now rather than at renewal: do you need 1040 tax prep software at all? Once e-file is solved, the federal calculation is free, the gathering already lives in your document tool, and what you'd be replacing is the review screen and the state returns, which is a smaller thing to rebuild than a tax engine, though not nothing. The vendors can't bundle their way out of that question, and the day the e-file application clears is the day it stops being hypothetical.
3. FreshBooks says its connector reaches every business on the login. For a tax practice, that's the feature.
On Friday, FreshBooks announced an Anthropic-verified connector (the plug that lets an AI assistant reach the data inside a piece of software): owners "can find FreshBooks in Claude's connector store and link their account in about 30 seconds, no setup required in FreshBooks." The release positions it as a way for owners who "don't have a financial background or work with outside bookkeepers and accountants" to "get answers to financial questions… without hiring for support or outside expertise." It's read-only today; FreshBooks' support article says Claude "can only read your live data to answer questions, not change data in your account," with invoice, expense, and journal-entry writes "coming soon." FreshBooks says 72% of its AI-using customers are already on Claude, which is 72% of the roughly 360 AI-using customers in a 600-customer survey with no fielding date, so hold it loosely.
The sentence that matters is in the support article. "Connecting Claude to your FreshBooks account will connect it to all active businesses associated with your login," and what Claude can see follows the user's role in FreshBooks. The Accountant role covers bank reconciliation, the chart of accounts, journal entries, invoices, payments, expenses, and bills.
FreshBooks' Accountant Hub lets an accountant reach every client business from one login: "once logged in, select your client's business from the list of businesses." Put those together and one connection from an accountant login reaches every client file the firm serves. That's my reading of FreshBooks' own documentation, and I haven't confirmed it with a test connection or with FreshBooks; if you run FreshBooks clients, the test takes a minute.
Compare the two connectors you're more likely to have. Xero's is read-only ("Claude has no write actions") and works one client organization at a time, which I covered last week. Google's QuickBooks connector in Gemini reaches nothing until a user signs in to a QuickBooks company through it, one company at a time. If my reading holds, FreshBooks, the smallest of the three in CAS, shipped the connector with the widest scope, and multi-client reach is the thing an accountant would have asked for first.
FreshBooks' customer is the reason the scope matters. It's used by small businesses that mostly do their own books, and those businesses come to a practice for tax rather than bookkeeping. The work a tax-only client generates is gathering and verification: pull the year's P&L, the general ledger, and the trial balance; check that the categorization holds up; find the owner draws that were coded as expenses; ask the questions before the return is prepared. That's the work a read-only connector across every client is built for, and it's the same gathering layer the OpenTax section says the return engine is about to sit underneath. The scope FreshBooks shipped is the scope the year-end gathering pass needs.
What this means for you. If you have FreshBooks clients, connect from an accountant login this week and see whether the client businesses appear. If they do, you have a year-end gathering pass that reaches every one of them from one place, and the first job is writing the questions you want asked of each file, because the connector will answer whatever it's asked. If you don't have FreshBooks clients, you now have the scope specification to put to Xero and Intuit: every client, one login, role-scoped, read-only until you say otherwise. Ask for it.
4. 87% of firms using AI have no written policy, and the profession just picked a starting point to test
Financial Cents surveyed 486 bookkeeping and accounting professionals across North America in July and August, and the sample is this brand's reader: mostly firms of two to 30 people, mostly owners and partners, most of them offering CAS work. It's a vendor survey with no stated recruitment method, and the closest comparison is another vendor's: Karbon's State of AI in Accounting 2026, fielded in October 2025 with a similar firm-size mix, found 21% of firms with an AI policy against Financial Cents' 13%, and more time saved where a policy exists. Two vendors, two samples that skew toward their own customers, and the same shape. Adoption is near-universal and shallow: on Financial Cents' own scale, 11% of firms describe themselves as running with AI, and two-thirds are walking or crawling. Firms with a written policy report measurable ROI at 35% against 14% for firms with none, and the running firms at 56% against the crawlers' 4%.
The same Tuesday, the AICPA launched the Council on AI Risk in Tax, led by former IRS Commissioner Danny Werfel, with the National Association of Enrolled Agents and the Federation of Tax Administrators named as participants, to "test, refine and strengthen an AI risk framework for tax." Werfel states: "Tax is high-stakes, and we cannot assume that emerging AI tools will police themselves." The framework is Werfel's own August 31 article in The Tax Adviser, and the half a firm should read isn't the one in the coverage. The 20-risk register is for tax authorities. The register for tax preparers is 15 risks in five categories, from hallucinated advice and client data leakage through failure of professional judgment, malpractice exposure, and client trust erosion, and its last category is the argument this roundup has made for two years: "Skill erosion: Staff develop expertise in operating AI tools rather than in the underlying tax law those tools are applying," "Vendor dependence," and "Institutional knowledge loss."
Three limits, so the register is read for what it is. It's Werfel's proposal rather than an AICPA standard; the council's job is to test it, and no deliverable dates were given. It's built for tax, so a bookkeeping practice borrows the structure rather than the list. And it doesn't mention fees.
What it does have is a priority scheme, "Tier 1 risks must be addressed before a deployment launches," and Werfel's operating rule: pilot with limited scope, test, then scale. That's a one-page firm policy with the headings already written, and the profession's body has agreed to test it. The audit regulator supplied the frame a week earlier, when PCAOB Chair Logothetis was reported saying "The science, the transaction part of it, can be replaced, can be done by AI… It's been done already," while skepticism and judgment "requires people."
What this means for you. First, place your firm on the Financial Cents scale without flattering it, because the ROI numbers say the difference between crawling and running is 4% against 56%, and the tools aren't what separates them. Second, if you don't have a written policy, don't start from a blank page. Take Werfel's 15 risks, cross out the ones that don't apply to your services, and write one sentence under each that says what your firm does about it. That's a policy, and it takes an afternoon.
Quick hits
Accrual shipped Arc, the third vendor in two weeks and the fourth this year to offer to hold your process and run it, and the third to say nothing about whether you can leave with it. Announced Tuesday and "available now to Accrual customers," Arc launches with 20 unnamed agents and lets firms "teach Arc their own processes, save them as reusable agents, share them across teams and schedule recurring work," with results delivered "alongside a record of the sources, tools and actions used to produce them." No price is stated, the only customer voice is Armanino's, and Accrual's CEO said on September 2 the company is "focused on serving Top 100 accounting firms," so this isn't a product you can buy. It matters as a pattern: Xero's XeroForce did this in May, and Intuit, FloQast, and now Accrual have each offered the same in the last two weeks; Xero's answer on export can be inferred from its architecture, since its agents live inside Xero, while the other three announcements say nothing about what happens to your process on exit. Before you hand any vendor a process, that clause is the one to settle.
Ramp moved from money-out to money-in, and your AR retainers just got a $0 comparator. Ramp Accounts Receivable, launched Tuesday for "U.S.-based, single-entity businesses with QuickBooks Online and NetSuite integrations," turns an uploaded contract into an invoice draft, prepares the overdue follow-up for someone to review and send, and matches incoming payments to open invoices. The FAQ is more precise than the marketing: the free tier covers invoicing, collections, and cash application, while revenue recognition and subscription billing sit in a paid add-on, and Ramp AR isn't sold on its own. Last week the spend platform shipped a ledger; this week it shipped receivables. If you invoice, chase, and apply cash for QBO clients on Ramp, know that number before the client does.
The largest firms have moved AI off the productivity line. EY US reported a 5% PCAOB findings rate for 2025, its best ever and down from 28%, and, as Accounting Today reported from EY's audit quality report, tied it to AI agents in the hands of "100% of EY US audit and technology risk professionals." KPMG US set up a venture studio "working like a start-up at the edge" to build companies that compete with its own service lines and, per Fortune, built an in-house citation checker after 40 of 45 citations in a June report were found to be fabricated. The most useful line came from PwC's Shawn Panson with no announcement attached: "The LLM is just the raw brain, but the harness is the scaffolding… we have our humans still checking everything," built so the firm can "ingest that data once and use it multiple times rather than having to ask the client for very similar pieces of information multiple times." Rented intelligence, an owned harness (the firm's own scaffolding around the model), human review, and ask the client once: that's the architecture, and it's being built for the largest firms first, while KPMG's edge companies are leaner versions of the firm that will compete with everyone below it. I don't read this as good news for a small firm: the parts got cheaper this week, but the tools that run a whole practice are going to the Top 100 first (see Accrual above), the firms that get them are spinning out lean competitors, and that's the four-way squeeze arriving from above, and the strongest case I've seen for small firms combining into larger ones with the resources to answer it.
What shipped in the stack, and which of it is already on in a client file. Copilot in Excel now links to where it changed a workbook and marks its own edits in the Show Changes pane, so Excel's own change log now names Copilot as the editor. Digits added eleven fixed-rule quality checks (negative balances, non-zero clearing accounts, multiple categories per vendor, PII in memos) that run as transactions post, free on public plans, with the line "Digits won't silently change the books." Xero renamed Files to Documents and added auto-matching of documents to transactions behind a toggle, which is Hubdoc's job built natively, and Anthropic folded connectors, plugins, and partners into one Claude Marketplace with "more than 2,000 available today." The question for each: is it switched on in a client file, and did anyone at the firm decide that?
The week in one line
This was the week the parts got cheaper and the subscription still went up. The published price of a token fell by a fifth to a half on Tuesday; the calculation engine at the center of the profession's highest-volume product was published for free the same day; and a connector that, on the vendor's own documentation, reaches every client on the login takes 30 seconds to set up. Xero's US price rises on October 1 and Intuit's rose in August, both with AI as the reason. The intelligence inside the platforms is deflating while the platforms inflate, and the gap between those two curves is either your margin or theirs, depending on who has done the arithmetic.
What didn't get cheaper is deciding. Financial Cents' survey says 87% of firms using AI have no written policy, and the firms that do report measurable returns at two and a half times the rate of firms without one. Every story this week ends at the same desk: which model runs which work at what cost per accepted result, whether a free engine gets a second-check pass, who at your firm can connect a client's books to what, and what happens when a vendor asks for your process. Those are policy decisions, and they don't wait for the platform to price them.
If those decisions are sitting on your desk unmade, book a free consultation at theaiaccountant.ai/consultation. It's a one-on-one call with me to look at where your firm stands with AI and which of these decisions to make first.

