The A to Z of AI Accounting Software (Atozai)

Can Claude close the books?

By Damon Anderson, 9 September 2026. 12 minute read.

Check yo self before you wreck yo self.
— Ice Cube

Last week I wrote about the Xero faux pas in paying an influencer to tell small businesses they could connect their ledger to Claude and stop paying their accountant £800 a month. That story was about a company's boardroom forgetting who helped build it. But underneath the hullabaloo was a question that was accidentally asked out loud, and one a very large number of accountants have likely been grappling with.

Can I just use Claude to do the work?

Not "will AI replace me." The honest answer to that one is that it already is replacing parts of the role. But "will it replace the IP one has spent their entire career building?" I think not. This is really about which parts go to AI tooling, which already have, and what is left standing for accounting professionals to keep putting food on the table.

First a quick level-set 101, because I'm often guilty of assuming everyone loves this stuff as much as I do. Claude is an AI assistant made by Anthropic, the same species of thing as ChatGPT or Gemini. Two years ago these products produced confident nonsense. Enough to be helpful, but not enough that you'd trust them with anything that mattered. Soufflé recipes, yes. Trusting due care has been satisfied, no. It's now a fairly trivial matter for Claude to read a year of bank transactions, spot the odd ones, explain why, and draft the client email about it. All in the time it takes for your Guinness to settle.

For the other end of that spectrum: Humanity's Last Exam is a benchmark written by experts specifically to be too hard for AI. At launch in early 2025 the best models scored single digits. Claude's Fable 5.1 now tops it at 65%. And it's only just got started.

Other frontier models are available, by the way. Claude gets named throughout because it is the category leader and the one most accountants I speak to are actively trying. Read it as shorthand for general-purpose intelligence.

So what exactly are the other 389 companies for? That is not a rhetorical number. It is how many accounting and finance vendors I track, and a great many of them sell AI that does essentially the same things Claude does.

Sometimes it can be hard to get a straight answer, because this conversation has gone the way politics on X (the artist formerly known as Twitter) has gone. Extreme views designed to get attention. One camp says the profession is finished. The other says it is all hype, probabilistic autocomplete dressed up by people selling something. Both have an interest. The end of accounting is a good story if you are raising money, and so is the debunking if you are booking speaking slots. Nobody has ever gone viral by saying it is complicated and depends on the use case. Certainly not any of Xero's influencers it would appear.

Which is a shame, because somewhere in the middle is where almost everyone actually is. I spend the majority of my allocated daily screen time on that question, so let me answer it properly.

So, can I just use Claude?

Yes. You can.

If you know what you are looking at, you can connect Claude to a ledger and get real work out of it today, in a fraction of the time it once took. I do it. I can do a hundred times more than I could a year ago, and so can any accountant who has sat down with these tools rather than reading about them from each side of the AI spectrum. Anyone telling you otherwise is either not using them or selling you something. It is also unrecognisably better than six months ago. Actually, I'd argue ten days ago with Fable 5.1. Any argument built on "the model cannot do X" might have a point, but that rhetoric has a bloody short shelf life.

The Excel test

You could run an entire accounting practice on Excel. There is essentially nothing in accounting it cannot be made to do, and people did for decades. Some still do.

It would also not be entirely advisable. Not because Excel is bad software, but because nobody built it around what regulators expect, what auditors ask for, what a partner needs to see before they sign. Every one of those has to be carried in someone's head and applied by hand, every time. And when it goes wrong, it goes wrong silently, in a cell reference nobody checked. Confident nonsense.

That is the shape of the Claude question. Not "is it powerful," but "what has it been built to guarantee."

Because Claude is a generalist that is very good at helping with accounting. It is not an accounting company. Nobody at Anthropic is sitting with a fifteen-partner firm mapping how a VAT return moves through a practice. That is not a criticism, it is the brief: be good at law and medicine and code all at once. Being domain-committed is the antithesis of that.

The software companies I track are doing the other thing. Forward-deployed engineers inside practices. Multinational tax codes mapped. Traceability built because they have sat across a table from someone who has to defend the number. That knowledge is in no training corpus. It lives in firms, and the only way in is to go and sit in them.

Probabilistic on top, deterministic underneath

A large language model is probabilistic. That word does the rounds and is doing a lot of work, so plainly: it predicts the most plausible next thing rather than calculating a correct answer. Ask it the same question twice and you can get two different replies, both reasonable, one possibly wrong.

That is a spectacularly cool thing when it comes to drafting, summarising and explaining. It is outright dangerous on its own for anything that has to be repeatable, consistent and defensible.

Accounting has both kinds of work in it. Explaining a variance is probabilistic work. Balancing a double entry is not, and a system that arrives at it by inference rather than by rule will eventually be confidently wrong in a way nobody catches.

Although watch where accountants actually resist and it does not split along that line. Extraction and reconciliation are rule-bound and the profession has handed them over happily for years. Tax is arguably more rule-bound still, and there the resistance is total. So the dividing line is not whether the work follows rules. It is whether the consequence of getting it wrong lands on you personally.

When you are the one who signs, you need the parts that can be guaranteed to actually be, well, guaranteed. The vendors worth paying attention to have built accordingly: the model on top doing the language-shaped work, and underneath a deterministic, canonical system of record doing the parts that must never be a guess.

Connect a general-purpose model directly to a ledger via an MCP connector and boom, job done. Count your savings and back to Instagram. Which is fine, right up until the point it isn't.

What the hell is an MCP anyway?

I have spent the last few weeks scoring how deeply every vendor in the A to Z plugs into external AI models, using something called MCP. 355 of 389 assessed and counting, and I am fairly sure it is the first time this has been measured across the accounting software market rather than taken from vendor marketing.

MCP stands for Model Context Protocol, which still tells you nothing. Another way to think about it is as a standard plug socket. Before it existed, connecting Claude to your ledger meant building a one-off bridge. Now a vendor publishes one socket and any assistant can plug in. Of the 355 I checked, 97 have one. Intuit shipped QuickBooks into both Claude and ChatGPT in July, able to create and send invoices from the chat window, not just read them.

What matters is not whether a vendor has a socket, but what you can do once plugged in. So the score runs 0 to 5, and it deliberately does not count tools, because tool counts are easy to inflate and impossible to verify from outside.

The MCP depth scale, 0 to 5, as shelves in a hardware shop

The line between 3 and 4 matters most in practice. Both can change your books. A tier 3 server hands an agent everything the account can reach, so an assistant you connected for one client can act across all of them. A tier 4 server can be fenced to a single entity or action. Same capability, very different consequences when something goes wrong.

Here is a slice:

MCP depth (0-5)
5
Audit trail included
4
Only what you need
2
Look, don't touch
Three shown per band. Across the 355 assessed: six vendors at 5, 22 at 4, 19 at 3, 24 at 2, 26 at 1, and 258 with nothing published.

Every score, and the reasoning behind it, is on the vendor's profile in the A to Z. Before you read it as a league table, do not. A low score is not a bad product, and emphatically not a bad AI product. Several of the sharpest AI companies here sit at zero on purpose, and I will come back to why.

Two things in that table point in opposite directions.

Not many household names at the top (I hear you say)

Six companies reached tier 5. Not one of them is an incumbent ledger.

Xero, QuickBooks, Sage Intacct and NetSuite all cap at four or below, failing on the same thing: no log of what the AI did. XBert writes every action to its audit log. Mighty tags anything an AI created into a dedicated "AI Created Records" view. Neither is a large company. Both have better AI governance than any platform on the board. An audit trail is the difference between an AI that helped and an AI you can answer questions about later. If a client asks in two years why a transaction was coded that way, "the assistant did it" is not an answer.

There is a technical reason they cap there. An MCP server is a wrapper: it exposes what the API underneath can already do, so its ceiling was set years earlier in the data architecture. You cannot bolt fine-grained AI governance onto a coarse legacy API, which is why a twenty-year-old platform with far more engineers loses a tier to a startup.

Two caveats. Several big names blocked my checks, so "no incumbent at tier 5" is partly "I could not check them all". And a zero means I found nothing published, not that nothing exists. If you're listed and I have you wrong, send me your MCP server URL and it goes into the next re-check.

Who is the intelligence, and who is the plumbing

Now the other direction, and it tells you what each company thinks it is.

Every vertical accounting agent I scored came back at zero. Ravical, Dytto, Combinely, Sumary, Mimo, Artifact AI, Briefcase. No MCP server found. Read that as a scoreboard and you would conclude the specialists are behind, which is absurd, since several are AI companies and nothing else.

What it reflects is which end of the socket you want to be on. Publishing a server says you are content to be the data layer, with the intelligence living elsewhere in a model run by a lab you have no relationship with. The vertical agents publish almost nothing and consume constantly. They are doing the reasoning themselves.

So it is a map of who has decided to be the intelligence and who has decided to be the plumbing. Xero at four is not Xero winning. It is Xero volunteering, very capably, to be the ledger underneath somebody else's mind.

Which is a more uncomfortable reading of last week's advert than the one I gave it. I called it a company forgetting who helped build it. The MCP data suggests the direction was already chosen, quietly, in engineering decisions taken long before any influencer was paid. The advert did not reveal a change of heart. It said out loud what the architecture had already settled.

Seen that way the market sorts into three groups.

The ones who are the data. The incumbent ledgers, opening up so an outside assistant can reach in. Xero, QuickBooks, Sage. A bet that owning the record is the durable position and the reasoning layer is somebody else's problem.

The ones who consume it. The vertical agents, doing what Claude does but pointed at one industry. Basis, Artifact AI, Archie, Combinely, Sumary, Mimo, Ravical and Dytto. Not cleverer than a frontier model, but aimed at one profession's workflows with domain knowledge underneath. Ravical is a good example, founded in 2025 by Silverfin co-founder Joris Van der Gucht, so not a general AI team adding accounting afterwards. Its model is three steps: signal the work, do the work, price it by its value. That last step is built to move firms off hourly billing and onto outcomes.

The ones trying to be both. Digits, Campfire, Rillet and, on its own announcements from CEO Reuben Steenkamp, soon Briefcase. They have built the ledger itself so it can be interrogated by AI natively. No bolt-on step, because there is nothing to bolt onto. The hardest position to build, and the strongest if they pull it off.

Why this is really about services, not software

Sequoia published a thesis this year that the next enormous company will be a software company masquerading as a services firm. For every dollar the world spends on software it spends roughly six on services, and the industry has spent fifteen years fighting over the dollar.

Which is why the technical argument matters commercially. If accounting is shifting from buying tools to buying outcomes, the question is not which software has the best features but who is prepared to be accountable for the result. Nobody sells Claude as an outcome. You buy it as a tool and you remain the one on the hook. A vendor selling a deterministic, traceable, tax-code-aware system is implicitly selling a share of that accountability, because they have engineered it so the answer is guaranteed rather than generated. You cannot underwrite a guess.

So, is Claude angling to eat your lunch?

Yes in parts. Plenty will use it, and they will save real money doing so.

But Claude will not carry your PI cover. Nor will ChatGPT. Neither will be struck off, and nobody goes to jail for a language model. When it is wrong, and eventually it will be, the person holding it is you.

That is why the accountant stays primary. Not the software. The judgement, and the willingness to stand behind it. What accountants deserve is technology that knows where these models cannot be trusted, and has done the unglamorous engineering to make the difference safe.

Plenty claim that now, and there will be more, because building software has never been cheaper. Most will not make it. The ones that do will not win on model quality, because everyone is renting the same models. They will win on understanding who they are building for: people who carry other people's financial lives and take real risk doing it.

They will embrace the community rather than market at it. They will be authentic about what they are and are not. They will articulate what they stand for, not just what they have built. They will have people who care about accountancy's best interests at heart.

That is product meeting market with both sides standing level. Rarer than it sounds, and the entire reason I spend my days sorting the companies who mean it from the ones who say it.

So before you cancel the accountant and hand the ledger to a chatbot, remember the words of the great philosopher Ice Cube.


The A to Z of AI Accounting Software tracks 389 vendors on what their AI actually does, evidence and all. MCP depth scores are now live on every profile. Sign up free.